Deploy vLLM in production
To deploy vLLM in production: install it on Linux (pip or the vllm/vllm-openai Docker image), run vllm serve with your model, set --gpu-memory-utilization and --max-model-len, enable --api-key, then put a reverse proxy in front. The server listens on port 8000 with an OpenAI-compatible API. It serves only one model at a time and targets the throughput of many simultaneous users.
vLLM is the inference server designed to share a GPU across many parallel requests. This guide covers memory sizing (the real issue), installation, startup, Docker and systemd, the settings that matter, throughput measurement, and security, with one important correction: the --api-key option protects only some routes.
#What vLLM does, and what it requires
vLLM is an open-source inference engine created at UC Berkeley’s Sky Computing Lab. Its central idea, PagedAttention, manages the attention key-value cache in pages, like an operating system’s virtual memory. According to the project’s initial announcement in 2023, existing systems wasted a large portion of their memory, and vLLM achieved up to 24 times the throughput of Hugging Face Transformers and up to 3.5 times that of TGI on tests from that period. These figures are old and specific to that benchmark: they indicate a direction, not what you will achieve with your model and GPU.
As for prerequisites, the current documentation requires Linux and Python 3.10 through 3.13; on Mac, there is a separate path, vLLM-Metal, built on MLX. The server exposes an OpenAI-compatible API, listens on port 8000 by default, and serves only one model at a time: for multiple models, you need multiple instances.
#When to choose vLLM over Ollama
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The difference isn't the ability to process requests in parallel, which Ollama also offers, but how memory is shared. According to Ollama's FAQ, parallel processing of a model multiplies the context size by the number of requests: a 2,000-token context with 4 parallel requests becomes an 8,000-token context in memory, reserved in advance. vLLM allocates its cache in blocks, on demand, and groups active requests into the same computations.
| Criterion | Ollama | vLLM |
|---|---|---|
| Concurrent users | 1 to a few; OLLAMA_NUM_PARALLEL controls parallelism | Dozens of simultaneous requests |
| Served models | Several, loaded and unloaded on demand | One per instance |
| Setup | An installation command | Python, CUDA, and settings to tune |
| Quantifications | GGUF, broad selection | Hub formats (AWQ, GPTQ, FP8); GGUF partially |
| Monitoring metrics | Not covered in detail in this guide | Documented /metrics endpoint |
| Typical use | Personal workstation, small team | Internal service or product |
Rule of thumb: if fewer than three people use the model at the same time, or if you want to switch models often, Ollama is enough. Beyond that, or with a single model served continuously, vLLM justifies its complexity. The comparison guide explains the choice in detail.
#When vLLM is a poor choice
vLLM offers nothing to a solo user on an 8 to 12 GB GPU: there isn’t enough memory for a shared cache, and Ollama or llama.cpp start more simply. It’s a poor fit if you switch between five models throughout the day, since you need to restart an instance for each model. On a Mac, the approach is different and relies on MLX. Finally, if you need a team chat interface rather than a high-load API, a Ollama stack with Open WebUI is a better fit and requires less operational work.
#Sizing memory: the calculation to do first
A vLLM server is sized by its key-value cache, not its weights. After loading the model, vLLM reserves a fraction of GPU memory—92% by default according to the current configuration code—and dedicates everything left to the cache. That remainder determines how many conversation tokens can coexist, and therefore how many concurrent users you can serve.
Take Qwen2.5-7B-Instruct, whose model card lists 7.61 billion parameters, 28 layers, and 4 key-value heads (grouped attention). The 16-bit weights take about 15.2 GB. The cache for one token is 2 (keys and values) × 28 layers × 4 heads × 128 dimensions × 2 bytes, or 57,344 bytes—about 56 KiB.
| GPU memory | 92% reserve | Remaining cache space | Cache tokens (upper bound) | Equivalent number of 4,096-token requests |
|---|---|---|---|---|
| 24 GB | 22.1 GB | 6.9 GB | about 120,000 | about 29 |
| 48 GB | 44.2 GB | 28.9 GB | approximately 500,000 | about 120 |
| 80 GB | 73.6 GB | 58.4 GB | approximately 1,000,000 | about 250 |
These limits are high: compute buffers and CUDA graphs consume some of the remaining capacity, and the model may have a different profile. The method remains valid for any model: check the layer and key-value head counts in the specifications, calculate the cost per token, and divide what remains. If the logs report preemptions, the documentation recommends increasing gpu_memory_utilization or reducing max_num_seqs.
Two levers expand the cache without changing your card: load a quantized version of the model, which frees up some of the weights, or limit --max-model-len, which avoids reserving space for contexts no one uses. The first lever may cost a little quality; the second costs nothing as long as your requests remain short.
#1. Installation
The documentation recommends uv, which automatically selects the right PyTorch version for your CUDA driver. For an AMD GPU, installation uses a dedicated index; for Intel, TPU, or Ascend, plugins are available. In production, the Docker image avoids CUDA version conflicts and updates with a simple tag change.
#2. Start the server
The vllm serve command replaces the old python -m vllm.entrypoints.openai.api_server invocation, which the current documentation no longer uses. On the first launch, the weights are downloaded from Hugging Face: plan for the disk space (about 15 GB for a 7B model in 16-bit). By default, the server applies the model repository’s generation_config.json file, and therefore the sampling parameters recommended by its publisher; --generation-config vllm restores vLLM’s default values.
#3. Docker and systemd
The official vllm/vllm-openai image is the safest option. Mount the Hugging Face cache so you don’t download the weights again, along with a volume for the compilation cache: otherwise, each new container starts with an empty cache and recompiles its model artifacts. Note that the image runs as root by default; the documentation describes running it as an unprivileged user (--user 2000:0).
#4. The parameters that matter
| Parameter | Role | Advice |
|---|---|---|
| --gpu-memory-utilization | Fraction of GPU memory reserved (0.92 by default) | Lower it if another process uses the GPU; raise it if the logs show preemptions |
| --max-model-len | Maximum supported context | As low as possible: each context token costs cache |
| --max-num-seqs | Maximum number of requests in a batch | Lower this if you run out of memory |
| --tensor-parallel-size | Distributes the model across multiple GPUs in a node | Only if the model does not fit on a GPU |
| --api-key | Requires a key for some routes | See the security section: insufficient on its own |
| --generation-config vllm | Ignore the model's generation_config.json | Use when the responses differ from what you expect |
One documentation principle: if the model fits on a single GPU, distribution is probably unnecessary; if it doesn’t fit but does fit on a node, use tensor parallelism with --tensor-parallel-size. Already-quantized models load directly from the Hub without any special option: the --quantization option is only for dynamic quantization.
#5. Measure throughput correctly
The vllm bench serve command sends requests to the server and reports throughput, time to first token (TTFT), and inter-token latency. The documentation specifies that these benchmarks are mainly used to evaluate functionality and detect regressions, and recommends GuideLLM for testing a production server.
#Deployment and operation
Once sizing is complete, deployment always follows the same sequence. It applies to a team of about twenty people querying the same model with 7 to 8 billion parameters on a 24 or 48 GB card.
- 01Choose the model and formatOne model per instance. Prefer a repository that is already quantized or in 16-bit, depending on available memory.
- 02Calculate the cacheApply the per-token calculation from the sizing section to set --max-model-len and --max-num-seqs.
- 03Start in DockerUse the official image with the Hugging Face cache mounted and a pinned version tag instead of latest, to prevent an update from changing the behavior.
- 04Add the proxyReverse proxy with a route allowlist, TLS, and rate limiting, followed by the API key as an additional measure.
- 05MeasureRun a load test with vllm bench serve using different seeds, and record TTFT and total throughput.
- 06MonitorConnect the /metrics endpoint collection to your monitoring tool.
The signals to watch for are those indicating a lack of cache: preemptions in the logs, rising TTFT, and lengthening queues. The documentation notes that preemption, whose default mode is recomputation, protects the service but degrades end-to-end latency. If it becomes frequent, increase gpu_memory_utilization, reduce the context, or limit the number of simultaneous requests. As a last resort, add a GPU and distribute the model with tensor parallelism.
#6. Security and exposure: --api-key isn’t enough
Contrary to what you often read, vLLM can verify an API key with --api-key or the VLLM_API_KEY variable. But the security documentation emphasizes that the key protects only routes under /v1, /v2, /inference, and /cohere. Other routes remain unauthenticated, including inference routes outside /v1, control routes such as /pause or /abort_requests, and /health. So never rely on --api-key alone.
- Reverse proxy
- Place nginx, Envoy, or a Kubernetes gateway in front of vLLM, with an allowlist of only the routes to expose, and block all others.
- Network
- A VPN or an isolated network: communication between the nodes of a distributed deployment is not secured by default.
- Development mode
- Never enable VLLM_SERVER_DEV_MODE=1 in production: it exposes dangerous routes.
- Limitations
- Apply rate limiting and request validation at the proxy level, as recommended by the documentation.
- Logs
- Log who sends what for debugging and auditing.
- Source: vLLM quickstart
- Source: vLLM security
- Source: vLLM with Docker
- Source: vLLM and PagedAttention announcement
Is vLLM better than Ollama in production?+
How do you run a vLLM server with an OpenAI-compatible API?+
How much VRAM does vLLM require?+
Is the --api-key option enough to secure vLLM?+
Does vLLM work on Mac or with an AMD card?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.