vLLM: what it is, who it is for, and when to use it ?
vLLM is an open-source inference engine, created in 2023 in Berkeley, that serves an LLM to multiple users simultaneously on GPUs, behind an OpenAI-compatible API. Thanks to PagedAttention and continuous batching, it delivers several times the total throughput of a traditional server as soon as requests arrive concurrently. It does not make a single conversation faster: for solo use on your own computer, Ollama or LM Studio remain the right choice.
vLLM is an open-source inference engine designed to serve an LLM to multiple users simultaneously, on GPUs, behind an OpenAI-compatible API. As of September 20, 2026, it is the standard choice for making an open-weights model available to a team or application. It is not a competitor to Ollama on your workstation: the two tools address different needs. This page explains what vLLM does, what it requires, and how to tell whether you need it.
#vLLM in three sentences
vLLM is a Python library and inference server created in 2023 at UC Berkeley’s Sky Computing Lab, released under the Apache 2.0 license, a permissive license that allows unrestricted commercial reuse. It loads a Hugging Face model, most often in safetensors format, on one or more GPUs and exposes it through an OpenAI-compatible HTTP API, allowing you to connect any client already written for the OpenAI API without changing a line of application code—only the base URL. Its purpose can be summed up in one word: throughput, meaning the total number of tokens produced per second by the GPU when ten, fifty, or two hundred different requests arrive at the same time and all need a fast response.
#The problem vLLM solves
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
When an LLM generates text, it keeps a record in memory of everything it has already read and written: the KV cache. This cache grows with each token, and its final size is unpredictable because you don't know in advance whether the response will be twenty words or two thousand. First-generation inference servers therefore reserved, for each request, a contiguous block of memory sized for the worst case.
The numerical result is described explicitly in vLLM's founding paper (Kwon et al., SOSP 2023) and reiterated by the project team itself: in existing systems, 60 to 80% of the memory reserved for the KV cache was simply wasted, due to both excessive precautionary reservation and fragmentation that leaves unusable gaps between allocated blocks. Less actually usable memory means mechanically fewer requests served in parallel on the same card, so a GPU costing several thousand euros ultimately runs at a negligible fraction of its theoretical capacity even though the electricity bill and hardware amortization remain unchanged.
#PagedAttention and continuous batching
vLLM is based on two ideas. The first, PagedAttention, is borrowed from operating systems: instead of one large contiguous block per request, the KV cache is split into small fixed-size pages, allocated on demand and stored anywhere in memory. A table maps the logical order of tokens to the physical location of the pages, exactly like a computer's virtual memory. Waste falls below 4%, according to the authors.
The second idea is continuous batching, which complements the first. A conventional server groups requests into batches and waits for every request in a batch to finish before starting another: the short request that could have resumed immediately waits unnecessarily behind the long one that still monopolizes the GPU. vLLM, by contrast, completely rebuilds the batch after every generated token, at each decoding step. As soon as a response finishes, its place in the batch is immediately given to a pending request in the queue, and the GPU never sits idle doing unnecessary padding. It is the combination of this constant recycling of slots and paged memory that produces most of the throughput gain observed over inference engines designed for only one user at a time.
- Shared prefixes
- When multiple requests start with exactly the same text (a system prompt, a shared document), the corresponding pages are computed only once and shared among them. This is prefix caching.
- Tensor parallelism
- A model too large to fit on a single card is automatically distributed across multiple GPUs in the same machine with a single startup option, --tensor-parallel-size.
- Server-side quantization
- vLLM reads AWQ, GPTQ, and FP8 formats, designed for batched GPU computation. GGUF is supported experimentally only.
- OpenAI-compatible API
- The /v1/chat/completions and /v1/completions routes respond exactly like OpenAI's: an existing application simply changes its base URL, without rewriting a single line of business logic.
#What vLLM loads into memory
This is the most common surprise for anyone coming from Ollama and expecting to reuse their habits directly. Ollama downloads a 4-bit quantized GGUF by default, designed to fit on a consumer graphics card. vLLM, by contrast, loads the weights published on Hugging Face as-is by default, most often in BF16 (16-bit precision), or roughly 2 GB per billion model parameters. The same model therefore takes roughly three times more memory under vLLM than under Ollama, even before counting the KV cache, which grows with each active conversation. This explains why a card sufficient for Ollama may turn out to be far too small for vLLM without changing the weight format.
| Model | BF16 (vLLM default) | GGUF Q4_K_M (default Ollama) | Minimum card in BF16 |
|---|---|---|---|
| Qwen 3 8B | 16 GB | 5 GB | RTX 4090 or 5090 (24-32 GB) |
| Gemma 4 12B | 24 GB | 7 GB | RTX 5090 (32 GB) |
| Qwen 3 14B | 28 GB | 9 GB | RTX 5090 (32 GB), short context |
| Mistral Small 3.2 24B | 48 GB | 14 GB | 2 × RTX 5090 (64 GB) |
| Qwen 3.8 27B | 54 GB | 16 GB | 80 GB card, or 2 × RTX 5090 with a short context |
| Llama 3.3 70B | 140 GB | 40 GB | 2 × 80 GB cards |
The most common workaround is to serve a version that is already quantized for GPU rather than the full original weights: a 4-bit AWQ or GPTQ model takes roughly the same amount of memory as an equivalent Q4 GGUF, and FP8 cuts BF16 size roughly in half on recent cards that support it natively. The vast majority of genuinely popular models have already been published in one of these formats on Hugging Face, often the same day as their official release.
#Supported hardware and systems
| Platform | Status | In practice |
|---|---|---|
| Linux + GPU NVIDIA (CUDA) | Primary target | The most thoroughly tested path. Compute capability 7.0 minimum, starting with Volta and Turing generations (RTX 20). |
| Linux + AMD GPU (ROCm) | Supported | Recent Instinct and Radeon cards. A dedicated Docker image is recommended. |
| Windows | No native version | Use WSL2 with a NVIDIA GPU. |
| Mac Apple Silicon | GPU supported since 22/09/2026 (separate plugin) | The official vllm-metal plugin, announced on September 22, 2026, brings vLLM's scheduler, KV-cache paging, and OpenAI-compatible server to Apple Silicon, with MLX and Metal for execution. It is installed separately from the main package and is still relatively new: verify your chip's compatibility before migrating a production workload. |
| CPU only (x86, ARM) | Supported | Useful for testing an integration, not for serving. |
#Get it running in five minutes
On a Linux machine with a NVIDIA GPU and up-to-date drivers, the complete installation fits in a few lines in a clean Python environment, with no complex prior configuration to write. The vllm serve command downloads the specified model from Hugging Face on the very first launch, loads it into memory, and immediately opens the API on port 8000 by default, ready to receive requests in the standard OpenAI format.
The --max-model-len option is worth setting explicitly from the very first launch: without it, vLLM sizes the KV cache for the maximum context window advertised by the model, sometimes 128,000 tokens or more on recent models, and simply refuses to start if the card's available memory cannot support that default sizing—an error that frequently occurs on the very first attempt. Proper production deployment, with a Docker container, call authentication, continuous monitoring, and gradual scaling, is covered in a separate, more detailed guide.
- Deploy vLLM in production: Docker, security, monitoring
- Official vLLM documentation
- The PagedAttention paper (Kwon et al., 2023)
- Source: the official announcement of vllm-metal (22/09/2026)
- Source: the project's official GitHub repository, vLLM
#Who it is for—and who it is not for
| Your situation | vLLM? | Why |
|---|---|---|
| You chat alone with a model on your PC | No | No speedup for a single request, and three times more VRAM in BF16. |
| You have a Mac | Maybe, recently | The vllm-metal plugin (22/09/2026) brings GPU support through MLX/Metal, but it is still very new. MLX and llama.cpp remain the proven choices as of 28/09/2026. |
| You have 8 to 12 GB of VRAM | Rarely | A GGUF Q4 with partial CPU offloading is more useful. |
| A team of 5 to 50 people shares a model | Yes | Continuous batching serves everyone on a single GPU. |
| An application calls the model in bursts | Yes | Queue, high throughput, standard API. |
| You process 10,000 documents in batches | Yes | This is the use case where the throughput gap is most pronounced. |
- llama.cpp vs vLLM vs ExLlama: three engines, three use cases
- llama.cpp: what is it, and should you leave Ollama?
- Calculate the VRAM required for your model
#FAQ
Is vLLM faster than Ollama?+
Is vLLM free?+
Does vLLM work on Windows or Mac?+
Can you use a GGUF file with vLLM?+
How many users can a GPU serve with vLLM?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.