Intermediate 11 minInference

vLLM: what it is, who it is for, and when to use it ?

Direct response

vLLM is an open-source inference engine, created in 2023 in Berkeley, that serves an LLM to multiple users simultaneously on GPUs, behind an OpenAI-compatible API. Thanks to PagedAttention and continuous batching, it delivers several times the total throughput of a traditional server as soon as requests arrive concurrently. It does not make a single conversation faster: for solo use on your own computer, Ollama or LM Studio remain the right choice.

vLLM is an open-source inference engine designed to serve an LLM to multiple users simultaneously, on GPUs, behind an OpenAI-compatible API. As of September 20, 2026, it is the standard choice for making an open-weights model available to a team or application. It is not a competitor to Ollama on your workstation: the two tools address different needs. This page explains what vLLM does, what it requires, and how to tell whether you need it.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#vLLM in three sentences

vLLM is a Python library and inference server created in 2023 at UC Berkeley’s Sky Computing Lab, released under the Apache 2.0 license, a permissive license that allows unrestricted commercial reuse. It loads a Hugging Face model, most often in safetensors format, on one or more GPUs and exposes it through an OpenAI-compatible HTTP API, allowing you to connect any client already written for the OpenAI API without changing a line of application code—only the base URL. Its purpose can be summed up in one word: throughput, meaning the total number of tokens produced per second by the GPU when ten, fifty, or two hundred different requests arrive at the same time and all need a fast response.

i
Key takeaway
Ollama and LM Studio optimize the experience for one person on their machine. vLLM optimizes the throughput of a GPU shared by many requests. If you're alone in front of your screen, vLLM won't make your responses faster.

#The problem vLLM solves

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

When an LLM generates text, it keeps a record in memory of everything it has already read and written: the KV cache. This cache grows with each token, and its final size is unpredictable because you don't know in advance whether the response will be twenty words or two thousand. First-generation inference servers therefore reserved, for each request, a contiguous block of memory sized for the worst case.

The numerical result is described explicitly in vLLM's founding paper (Kwon et al., SOSP 2023) and reiterated by the project team itself: in existing systems, 60 to 80% of the memory reserved for the KV cache was simply wasted, due to both excessive precautionary reservation and fragmentation that leaves unusable gaps between allocated blocks. Less actually usable memory means mechanically fewer requests served in parallel on the same card, so a GPU costing several thousand euros ultimately runs at a negligible fraction of its theoretical capacity even though the electricity bill and hardware amortization remain unchanged.

#PagedAttention and continuous batching

vLLM is based on two ideas. The first, PagedAttention, is borrowed from operating systems: instead of one large contiguous block per request, the KV cache is split into small fixed-size pages, allocated on demand and stored anywhere in memory. A table maps the logical order of tokens to the physical location of the pages, exactly like a computer's virtual memory. Waste falls below 4%, according to the authors.

The second idea is continuous batching, which complements the first. A conventional server groups requests into batches and waits for every request in a batch to finish before starting another: the short request that could have resumed immediately waits unnecessarily behind the long one that still monopolizes the GPU. vLLM, by contrast, completely rebuilds the batch after every generated token, at each decoding step. As soon as a response finishes, its place in the batch is immediately given to a pending request in the queue, and the GPU never sits idle doing unnecessary padding. It is the combination of this constant recycling of slots and paged memory that produces most of the throughput gain observed over inference engines designed for only one user at a time.

Shared prefixes
When multiple requests start with exactly the same text (a system prompt, a shared document), the corresponding pages are computed only once and shared among them. This is prefix caching.
Tensor parallelism
A model too large to fit on a single card is automatically distributed across multiple GPUs in the same machine with a single startup option, --tensor-parallel-size.
Server-side quantization
vLLM reads AWQ, GPTQ, and FP8 formats, designed for batched GPU computation. GGUF is supported experimentally only.
OpenAI-compatible API
The /v1/chat/completions and /v1/completions routes respond exactly like OpenAI's: an existing application simply changes its base URL, without rewriting a single line of business logic.

#What vLLM loads into memory

This is the most common surprise for anyone coming from Ollama and expecting to reuse their habits directly. Ollama downloads a 4-bit quantized GGUF by default, designed to fit on a consumer graphics card. vLLM, by contrast, loads the weights published on Hugging Face as-is by default, most often in BF16 (16-bit precision), or roughly 2 GB per billion model parameters. The same model therefore takes roughly three times more memory under vLLM than under Ollama, even before counting the KV cache, which grows with each active conversation. This explains why a card sufficient for Ollama may turn out to be far too small for vLLM without changing the weight format.

Weights only, excluding the KV cache · calculated from the QuelLLM catalog · 20/09/2026
ModelBF16 (vLLM default)GGUF Q4_K_M (default Ollama)Minimum card in BF16
Qwen 3 8B16 GB5 GBRTX 4090 or 5090 (24-32 GB)
Gemma 4 12B24 GB7 GBRTX 5090 (32 GB)
Qwen 3 14B28 GB9 GBRTX 5090 (32 GB), short context
Mistral Small 3.2 24B48 GB14 GB2 × RTX 5090 (64 GB)
Qwen 3.8 27B54 GB16 GB80 GB card, or 2 × RTX 5090 with a short context
Llama 3.3 70B140 GB40 GB2 × 80 GB cards

The most common workaround is to serve a version that is already quantized for GPU rather than the full original weights: a 4-bit AWQ or GPTQ model takes roughly the same amount of memory as an equivalent Q4 GGUF, and FP8 cuts BF16 size roughly in half on recent cards that support it natively. The vast majority of genuinely popular models have already been published in one of these formats on Hugging Face, often the same day as their official release.

!
“vLLM used all my VRAM”
This is entirely intentional, not a bug or a memory leak. At startup, vLLM reserves 90% of total GPU memory by default, through the --gpu-memory-utilization option set to 0.9 by default: the model weights first, then the rest of that space for KV-cache pages. The more pages available, the more requests vLLM can serve in parallel without rejecting them. On a shared machine with other workloads, a video game, or another service, lower this value accordingly.

#Supported hardware and systems

Support status as of 20/09/2026, according to the project documentation
PlatformStatusIn practice
Linux + GPU NVIDIA (CUDA)Primary targetThe most thoroughly tested path. Compute capability 7.0 minimum, starting with Volta and Turing generations (RTX 20).
Linux + AMD GPU (ROCm)SupportedRecent Instinct and Radeon cards. A dedicated Docker image is recommended.
WindowsNo native versionUse WSL2 with a NVIDIA GPU.
Mac Apple SiliconGPU supported since 22/09/2026 (separate plugin)The official vllm-metal plugin, announced on September 22, 2026, brings vLLM's scheduler, KV-cache paging, and OpenAI-compatible server to Apple Silicon, with MLX and Metal for execution. It is installed separately from the main package and is still relatively new: verify your chip's compatibility before migrating a production workload.
CPU only (x86, ARM)SupportedUseful for testing an integration, not for serving.

#Get it running in five minutes

On a Linux machine with a NVIDIA GPU and up-to-date drivers, the complete installation fits in a few lines in a clean Python environment, with no complex prior configuration to write. The vllm serve command downloads the specified model from Hugging Face on the very first launch, loads it into memory, and immediately opens the API on port 8000 by default, ready to receive requests in the standard OpenAI format.

Install and serve a model
python -m venv vllm-env && source vllm-env/bin/activate
pip install vllm

vllm serve Qwen/Qwen3-8B --max-model-len 8192
Query an API like OpenAI's
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen/Qwen3-8B", "messages": [{"role": "user", "content": "Bonjour"}]}'

The --max-model-len option is worth setting explicitly from the very first launch: without it, vLLM sizes the KV cache for the maximum context window advertised by the model, sometimes 128,000 tokens or more on recent models, and simply refuses to start if the card's available memory cannot support that default sizing—an error that frequently occurs on the very first attempt. Proper production deployment, with a Docker container, call authentication, continuous monitoring, and gradual scaling, is covered in a separate, more detailed guide.

#Who it is for—and who it is not for

Your situationvLLM?Why
You chat alone with a model on your PCNoNo speedup for a single request, and three times more VRAM in BF16.
You have a MacMaybe, recentlyThe vllm-metal plugin (22/09/2026) brings GPU support through MLX/Metal, but it is still very new. MLX and llama.cpp remain the proven choices as of 28/09/2026.
You have 8 to 12 GB of VRAMRarelyA GGUF Q4 with partial CPU offloading is more useful.
A team of 5 to 50 people shares a modelYesContinuous batching serves everyone on a single GPU.
An application calls the model in burstsYesQueue, high throughput, standard API.
You process 10,000 documents in batchesYesThis is the use case where the throughput gap is most pronounced.

#FAQ

Is vLLM faster than Ollama?+
For a single user, no: the generation speed of an isolated request depends primarily on the GPU’s memory bandwidth, which is identical in both cases. The difference appears under concurrency. When multiple requests arrive at the same time, vLLM processes them in the same batch, and its total throughput far exceeds that of a server designed for one user.
Is vLLM free?+
Yes, entirely. The project is open source under the Apache 2.0 license, a permissive license that allows commercial enterprise use without fees or royalties of any kind, unlike some model licenses. The only real cost is the hardware you already own or renting a GPU from a cloud provider to run the server.
Does vLLM work on Windows or Mac?+
Not natively on Windows: the official project has no Windows version or public roadmap for one, so you have to use WSL2 with a NVIDIA GPU. A few community forks exist, but they remain unofficial. On Mac, an official plugin named vllm-metal, announced on September 22, 2026, finally provides GPU acceleration through MLX and Metal; before that date, only the CPU was usable, which removed any advantage of vLLM over Ollama or MLX on that platform.
Can you use a GGUF file with vLLM?+
The support exists technically but remains experimental and performs significantly worse than the native path. vLLM was designed from the ground up for Hugging Face weights in safetensors format and for quantizations intended for batched GPU computation: AWQ, GPTQ, and FP8. If you absolutely need GGUF format, llama.cpp and its llama-server remain the natural and best-optimized tools for this specific format.
How many users can a GPU serve with vLLM?+
It depends on the memory left for the KV cache after the weights are loaded and on the length of the conversations. With a quantized 8B model on a 24 GB card, several dozen short simultaneous conversations are realistic. The only reliable answer comes from a load test with your own prompts.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.