Advanced 11 minvLLM

Deploy vLLM in production

Direct response

To deploy vLLM in production: install it on Linux (pip or the vllm/vllm-openai Docker image), run vllm serve with your model, set --gpu-memory-utilization and --max-model-len, enable --api-key, then put a reverse proxy in front. The server listens on port 8000 with an OpenAI-compatible API. It serves only one model at a time and targets the throughput of many simultaneous users.

vLLM is the inference server designed to share a GPU across many parallel requests. This guide covers memory sizing (the real issue), installation, startup, Docker and systemd, the settings that matter, throughput measurement, and security, with one important correction: the --api-key option protects only some routes.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What vLLM does, and what it requires

vLLM is an open-source inference engine created at UC Berkeley’s Sky Computing Lab. Its central idea, PagedAttention, manages the attention key-value cache in pages, like an operating system’s virtual memory. According to the project’s initial announcement in 2023, existing systems wasted a large portion of their memory, and vLLM achieved up to 24 times the throughput of Hugging Face Transformers and up to 3.5 times that of TGI on tests from that period. These figures are old and specific to that benchmark: they indicate a direction, not what you will achieve with your model and GPU.

As for prerequisites, the current documentation requires Linux and Python 3.10 through 3.13; on Mac, there is a separate path, vLLM-Metal, built on MLX. The server exposes an OpenAI-compatible API, listens on port 8000 by default, and serves only one model at a time: for multiple models, you need multiple instances.

#When to choose vLLM over Ollama

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The difference isn't the ability to process requests in parallel, which Ollama also offers, but how memory is shared. According to Ollama's FAQ, parallel processing of a model multiplies the context size by the number of requests: a 2,000-token context with 4 parallel requests becomes an 8,000-token context in memory, reserved in advance. vLLM allocates its cache in blocks, on demand, and groups active requests into the same computations.

Ollama or vLLM: decision criteria
CriterionOllamavLLM
Concurrent users1 to a few; OLLAMA_NUM_PARALLEL controls parallelismDozens of simultaneous requests
Served modelsSeveral, loaded and unloaded on demandOne per instance
SetupAn installation commandPython, CUDA, and settings to tune
QuantificationsGGUF, broad selectionHub formats (AWQ, GPTQ, FP8); GGUF partially
Monitoring metricsNot covered in detail in this guideDocumented /metrics endpoint
Typical usePersonal workstation, small teamInternal service or product

Rule of thumb: if fewer than three people use the model at the same time, or if you want to switch models often, Ollama is enough. Beyond that, or with a single model served continuously, vLLM justifies its complexity. The comparison guide explains the choice in detail.

#When vLLM is a poor choice

vLLM offers nothing to a solo user on an 8 to 12 GB GPU: there isn’t enough memory for a shared cache, and Ollama or llama.cpp start more simply. It’s a poor fit if you switch between five models throughout the day, since you need to restart an instance for each model. On a Mac, the approach is different and relies on MLX. Finally, if you need a team chat interface rather than a high-load API, a Ollama stack with Open WebUI is a better fit and requires less operational work.

#Sizing memory: the calculation to do first

A vLLM server is sized by its key-value cache, not its weights. After loading the model, vLLM reserves a fraction of GPU memory—92% by default according to the current configuration code—and dedicates everything left to the cache. That remainder determines how many conversation tokens can coexist, and therefore how many concurrent users you can serve.

Take Qwen2.5-7B-Instruct, whose model card lists 7.61 billion parameters, 28 layers, and 4 key-value heads (grouped attention). The 16-bit weights take about 15.2 GB. The cache for one token is 2 (keys and values) × 28 layers × 4 heads × 128 dimensions × 2 bytes, or 57,344 bytes—about 56 KiB.

Available cache depending on the GPU (Qwen2.5-7B at 16-bit, 92% of memory, before compute buffers)
GPU memory92% reserveRemaining cache spaceCache tokens (upper bound)Equivalent number of 4,096-token requests
24 GB22.1 GB6.9 GBabout 120,000about 29
48 GB44.2 GB28.9 GBapproximately 500,000about 120
80 GB73.6 GB58.4 GBapproximately 1,000,000about 250

These limits are high: compute buffers and CUDA graphs consume some of the remaining capacity, and the model may have a different profile. The method remains valid for any model: check the layer and key-value head counts in the specifications, calculate the cost per token, and divide what remains. If the logs report preemptions, the documentation recommends increasing gpu_memory_utilization or reducing max_num_seqs.

Two levers expand the cache without changing your card: load a quantized version of the model, which frees up some of the weights, or limit --max-model-len, which avoids reserving space for contexts no one uses. The first lever may cost a little quality; the second costs nothing as long as your requests remain short.

→
A 7B in 16-bit on 24 GB serves about thirty 4,000-token conversations
This calculation explains why vLLM shines on 48 or 80 GB cards: cache headroom, not the speed of a single user, is what makes the difference. On a 12 GB card, the same model leaves almost no room for the cache.

#1. Installation

Documentation-recommended installation (NVIDIA CUDA)
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

The documentation recommends uv, which automatically selects the right PyTorch version for your CUDA driver. For an AMD GPU, installation uses a dedicated index; for Intel, TPU, or Ascend, plugins are available. In production, the Docker image avoids CUDA version conflicts and updates with a simple tag change.

#2. Start the server

Getting started with vllm serve
vllm serve Qwen/Qwen2.5-7B-Instruct \
  --host 0.0.0.0 \
  --port 8000 \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192 \
  --api-key "$VLLM_API_KEY"

The vllm serve command replaces the old python -m vllm.entrypoints.openai.api_server invocation, which the current documentation no longer uses. On the first launch, the weights are downloaded from Hugging Face: plan for the disk space (about 15 GB for a 7B model in 16-bit). By default, the server applies the model repository’s generation_config.json file, and therefore the sampling parameters recommended by its publisher; --generation-config vllm restores vLLM’s default values.

Check the server
curl http://localhost:8000/v1/models \
  -H "Authorization: Bearer $VLLM_API_KEY"

#3. Docker and systemd

The official vllm/vllm-openai image is the safest option. Mount the Hugging Face cache so you don’t download the weights again, along with a volume for the compilation cache: otherwise, each new container starts with an empty cache and recompiles its model artifacts. Note that the image runs as root by default; the documentation describes running it as an unprivileged user (--user 2000:0).

Container with mounted cache
docker run --rm --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v vllm-cache:/root/.cache/vllm \
  -p 8000:8000 \
  --ipc=host \
  -e VLLM_API_KEY=$VLLM_API_KEY \
  vllm/vllm-openai:latest \
  Qwen/Qwen2.5-7B-Instruct
systemd unit (Docker-free installation)
[Unit]
Description=vLLM OpenAI API
After=network.target

[Service]
Type=simple
User=vllm
EnvironmentFile=/etc/vllm/env
ExecStart=/opt/vllm/bin/vllm serve Qwen/Qwen2.5-7B-Instruct --port 8000
Restart=always

[Install]
WantedBy=multi-user.target

#4. The parameters that matter

Key vllm serve parameters
ParameterRoleAdvice
--gpu-memory-utilizationFraction of GPU memory reserved (0.92 by default)Lower it if another process uses the GPU; raise it if the logs show preemptions
--max-model-lenMaximum supported contextAs low as possible: each context token costs cache
--max-num-seqsMaximum number of requests in a batchLower this if you run out of memory
--tensor-parallel-sizeDistributes the model across multiple GPUs in a nodeOnly if the model does not fit on a GPU
--api-keyRequires a key for some routesSee the security section: insufficient on its own
--generation-config vllmIgnore the model's generation_config.jsonUse when the responses differ from what you expect

One documentation principle: if the model fits on a single GPU, distribution is probably unnecessary; if it doesn’t fit but does fit on a node, use tensor parallelism with --tensor-parallel-size. Already-quantized models load directly from the Hub without any special option: the --quantization option is only for dynamic quantization.

#5. Measure throughput correctly

The vllm bench serve command sends requests to the server and reports throughput, time to first token (TTFT), and inter-token latency. The documentation specifies that these benchmarks are mainly used to evaluate functionality and detect regressions, and recommends GuideLLM for testing a production server.

Load testing
vllm bench serve \
  --backend vllm \
  --model Qwen/Qwen2.5-7B-Instruct \
  --endpoint /v1/completions \
  --dataset-name sharegpt \
  --dataset-path CHEMIN/ShareGPT_V3_unfiltered_cleaned_split.json \
  --num-prompts 200
!
Repeating a benchmark inflates throughput
The documentation warns that rerunning vllm bench serve on the same server may reuse prompts left in the prefix cache and inflate the results. Change the seed with --seed or restart the server between measurements.

#Deployment and operation

Once sizing is complete, deployment always follows the same sequence. It applies to a team of about twenty people querying the same model with 7 to 8 billion parameters on a 24 or 48 GB card.

  1. 01
    Choose the model and format
    One model per instance. Prefer a repository that is already quantized or in 16-bit, depending on available memory.
  2. 02
    Calculate the cache
    Apply the per-token calculation from the sizing section to set --max-model-len and --max-num-seqs.
  3. 03
    Start in Docker
    Use the official image with the Hugging Face cache mounted and a pinned version tag instead of latest, to prevent an update from changing the behavior.
  4. 04
    Add the proxy
    Reverse proxy with a route allowlist, TLS, and rate limiting, followed by the API key as an additional measure.
  5. 05
    Measure
    Run a load test with vllm bench serve using different seeds, and record TTFT and total throughput.
  6. 06
    Monitor
    Connect the /metrics endpoint collection to your monitoring tool.

The signals to watch for are those indicating a lack of cache: preemptions in the logs, rising TTFT, and lengthening queues. The documentation notes that preemption, whose default mode is recomputation, protects the service but degrades end-to-end latency. If it becomes frequent, increase gpu_memory_utilization, reduce the context, or limit the number of simultaneous requests. As a last resort, add a GPU and distribute the model with tensor parallelism.

#6. Security and exposure: --api-key isn’t enough

Contrary to what you often read, vLLM can verify an API key with --api-key or the VLLM_API_KEY variable. But the security documentation emphasizes that the key protects only routes under /v1, /v2, /inference, and /cohere. Other routes remain unauthenticated, including inference routes outside /v1, control routes such as /pause or /abort_requests, and /health. So never rely on --api-key alone.

Reverse proxy
Place nginx, Envoy, or a Kubernetes gateway in front of vLLM, with an allowlist of only the routes to expose, and block all others.
Network
A VPN or an isolated network: communication between the nodes of a distributed deployment is not secured by default.
Development mode
Never enable VLLM_SERVER_DEV_MODE=1 in production: it exposes dangerous routes.
Limitations
Apply rate limiting and request validation at the proxy level, as recommended by the documentation.
Logs
Log who sends what for debugging and auditing.
FAQ
Is vLLM better than Ollama in production?+
It performs better when several users query the same model at once: it shares the cache by blocks and batches requests. Ollama remains simpler for a personal workstation or a small team and lets you switch models on the fly. vLLM serves only one model per instance.
How do you run a vLLM server with an OpenAI-compatible API?+
Using the vllm serve command followed by the model name. The server listens by default on http://localhost:8000 and provides OpenAI-compatible routes, including /v1/models and /v1/chat/completions. Specify --host and --port to expose it to the network, and add --api-key and a reverse proxy before exposing it.
How much VRAM does vLLM require?+
Enough for the model weights plus the key-value cache for your concurrent users. A 7B model in 16-bit weighs about 15 GB; on 24 GB with 92% reserved, about 7 GB remains for the cache—roughly thirty 4,000-token conversations. On 48 GB, about four times as many conversations.
Is the --api-key option enough to secure vLLM?+
No. It only protects the routes under /v1, /v2, /inference, and /cohere; routes such as /health, /invocations, and /pause remain accessible without a key. The documentation recommends placing a reverse proxy that allows only the desired routes, and never exposing the server directly to the Internet.
Does vLLM work on Mac or with an AMD card?+
Yes, with caveats. The documentation supports AMD GPUs through ROCm, Intel, and other accelerators. On Mac, it points to vLLM-Metal, which relies on MLX rather than PyTorch and requires models in MLX format. The main path remains Linux with a NVIDIA GPU.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.