llama.cpp vs vLLM vs Exllama
llama.cpp is the portable engine for personal use (GGUF, CPU, Mac, GPUs of all brands) and for Ollama and LM Studio. vLLM is designed to serve many users simultaneously on GPUs, with continuous batching; Red Hat measures up to 793 tokens/s versus 41 for Ollama on an A100. ExLlamaV2 is archived: its development continues in ExLlamaV3.
Under Ollama, LM Studio, or Jan is an inference engine, and that engine determines the accepted formats, usable hardware, and behavior under load. This guide compares llama.cpp, vLLM, ExLlama, and SGLang using criteria that can be verified in their repositories, corrects misconceptions, and provides a selection rule. It publishes no in-house throughput figures: only attributed third-party measurements appear here.
#An inference engine: what really changes between them
An inference engine transforms input tokens into output tokens. The applications you install (Ollama, LM Studio, Jan) include one; production servers (vLLM, SGLang) are one. Three differences affect how you use them: the model formats they accept, the hardware they can use, and how they handle multiple requests at the same time.
| Engine | License | Announced formats and hardware | Typical use |
|---|---|---|---|
| llama.cpp | MIT | GGUF; CPU, Apple Silicon, NVIDIA (CUDA), AMD (HIP), Vulkan, SYCL, and others | Personal workstation, laptop, lightweight server |
| vLLM | Apache 2.0 | Hugging Face models; FP8, INT4, GPTQ, AWQ, GGUF (experimental); NVIDIA GPU, AMD, Intel, and x86/ARM CPUs | Serve many users, production |
| SGLang | Apache 2.0 | Hugging Face models; FP4, FP8, INT4, AWQ, GPTQ | High-throughput server, shared prefixes |
| ExLlamaV3 | MIT | EXL3; consumer GPU NVIDIA | Latency on a personal GPU with TabbyAPI |
| MLX LM | MIT | Apple Silicon only | Mac: generation and fine-tuning |
#The model format determines the possible engine
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Before comparing speeds, check the format in which the model you want is published. A GGUF file loads in llama.cpp and therefore in Ollama, LM Studio, and Jan. Hugging Face weights in FP16 or FP8, or quantized in AWQ or GPTQ, load in vLLM or SGLang. A model converted for MLX loads with MLX LM on Mac, and an EXL3 model with ExLlamaV3. Switching engines often means downloading the model again in another format.
| Format | Natural engine | Another possible engine |
|---|---|---|
| GGUF | llama.cpp (and therefore Ollama, LM Studio, Jan) | vLLM, in experimental mode |
| Hugging Face weights (FP16, FP8) | vLLM, SGLang | MLX LM after conversion, on Mac |
| AWQ, GPTQ | vLLM, SGLang | Check depending on the engine |
| EXL3 | ExLlamaV3 | None |
| MLX | MLX LM | LM Studio and Jan, which announce MLX |
#llama.cpp: the portable standard
llama.cpp is a dependency-free C/C++ implementation whose stated goal is inference with minimal configuration across a wide range of hardware. Apple Silicon is prioritized (NEON, Accelerate, Metal), x86 CPUs are supported with AVX, AVX2, AVX512, and AMX, and NVIDIA (CUDA), AMD (HIP), Vulkan, and SYCL GPUs are supported. It offers 1.5- to 8-bit quantization and hybrid CPU/GPU inference, allowing you to run a model larger than the available VRAM.
A misconception to correct: llama.cpp is not limited to one request at a time. Its server advertises multi-user parallel generation and continuous batching, enabled by default, with slots configured by the --parallel parameter. It therefore serves a small team effectively; what it does not aim to match are vLLM’s large-scale production features.
- Strengths
- Portability (Mac, Linux, Windows, Android), ubiquitous GGUF formats, CPU and GPU hybrid operation, low dependency.
- Limitations
- No distributed deployment features comparable to those in vLLM; settings (context, slots, layers on the GPU) to know.
- Ecosystem
- Foundation for Ollama and LM Studio, also cited by Jan.
For compilation with CUDA, Metal, or Vulkan, see the compilation guides; the complete guide to llama.cpp details how to use it.
#vLLM: the server for many users
vLLM is an inference and serving library created at the Sky Computing Lab at the University of California, Berkeley. It is built on PagedAttention, which manages attention key and value memory in pages, and continuous batching, complemented by chunked prefill and prefix caching. It advertises FP8, INT4, GPTQ, AWQ, and GGUF formats, an OpenAI-compatible API, and support for NVIDIA, AMD, and Intel GPUs, as well as x86, ARM, and PowerPC CPUs.
Two misconceptions to correct. First, vLLM is no longer limited to NVIDIA or CPU-less setups: its repository indicates support for multiple GPUs and CPUs. Second, it is no longer without GGUF support: its documentation describes that support as highly experimental and poorly optimized, useful mainly for reducing memory usage. For everyday GGUF use, llama.cpp remains the standard approach.
- Strengths
- Throughput under load, page-managed cache memory, OpenAI-compatible API, parallelism (tensor, pipeline, expert), and numerous formats.
- Limitations
- Heavier to install and configure than llama.cpp; designed for GPU servers; GGUF is not its domain.
- Ecosystem
- The default choice when multiple people or applications query the same model at the same time.
The guide on vLLM explains what the engine is; the production deployment guide covers configuration and monitoring.
#ExLlama: V2 archived, V3 in development
The ExLlamaV2 repository notes that the project is currently archived and development continues in ExLlamaV3. Many comparisons, including the previous version of this page, still present V2 as the state-of-the-art option. The ExLlamaV3 repository announces the EXL3 quantization format, tensor and expert parallel inference, CPU offloading for mixture-of-experts models, continuous batching, speculative decoding, and an OpenAI-compatible API through TabbyAPI, its recommended server.
ExLlamaV3 targets consumer GPUs, not production servers or Macs. If you have a NVIDIA card and want the best balance between quality and size, EXL3 quantization is worth trying; first check that your model appears in the repository’s list of supported architectures.
#SGLang, MLX LM, and others
- SGLang
- A serving framework described as targeting low latency and high throughput, from one GPU to large clusters. It advertises RadixAttention for prefix caching, continuous batching, PagedAttention, and speculative decoding. It competes with vLLM on servers; the dedicated guide covers it in detail.
- MLX LM
- Python package for generating text and fine-tuning models on Apple Silicon with MLX. It doesn't work outside of Mac. The MLX versus llama.cpp guide compares the two on Mac.
- TensorRT-LLM
- NVIDIA engine, highly performant on its cards but more demanding to set up; consider it only for a production NVIDIA fleet.
#What the published measurements say
Throughput depends on the hardware, model, quantization, query length, and engine version, and changes every month. This guide therefore presents no in-house measurements and advises you to be wary of tokens-per-second tables without a stated protocol. One third-party source does publish a complete protocol: Red Hat, in August 2025.
| Item | Value published by Red Hat |
|---|---|
| Hardware | One NVIDIA A100-PCIE-40GB card |
| Model | Llama 3.1 8B Instruct (FP16 on the Ollama side) |
| Versions | vLLM 0.9.1; Ollama 0.9.2 |
| Test tool | GuideLLM 0.2.1, from 1 to 256 concurrent users |
| Maximum throughput | 793 tokens/s for vLLM versus 41 for Ollama |
| Peak P99 latency | 80 ms for vLLM versus 673 ms for Ollama |
Read with reservations: the article is associated with Red Hat AI products, so it comes from a commercial player in the field; the tests cover Ollama rather than llama.cpp directly, using default settings; and they date back several versions. The robust takeaway is qualitative: under heavy concurrency, a serving engine with aggressive batching outperforms an application designed for a single user. For solo use, throughput is not what will hold you back.
#Memory: what each engine reserves
A model’s memory consists of its weights, context cache (KV), and headroom. The weights are easy to calculate: an 8-billion-parameter model uses approximately 16 GB in FP16 (8 billion times 2 bytes) and approximately 5 GB in Q4, according to the site’s rule of thumb. The context cache grows with conversation length and the number of simultaneous requests.
| Engine | Documented behavior | Practical consequence |
|---|---|---|
| llama.cpp | User-set context; parallel slots with a unified cache | You set the context size and number of slots |
| vLLM | Pre-allocates part of GPU memory for the cache, 92% by default | On a 24 GB card, approximately 22 GB are reserved immediately |
| ExLlamaV3 | 2- to 8-bit cache quantization | The cache can be compressed to fit in VRAM |
The key point: vLLM reserves memory from the start, making it effective for serving but poorly suited to a card shared with other applications. Reduce gpu_memory_utilization if the card is also used for something else.
#Which engine for which use case
| Your situation | Recommended engine | Reason |
|---|---|---|
| Personal use on Mac, PC, or laptop | llama.cpp, via Ollama or LM Studio | Portable and simple |
| Mac Apple Silicon, throughput-focused | MLX LM or llama.cpp | MLX is designed for Apple Silicon; compare on your models |
| A team or application queries the model | vLLM or SGLang | Continuous batching, cache pages |
| A consumer-grade NVIDIA card, maximum quality | ExLlamaV3 with TabbyAPI | EXL3 quantization |
| Mixed hardware, CPU, AMD, Intel | llama.cpp | Broad hardware support |
| GGUF-only model | llama.cpp | vLLM supports it only experimentally |
The engines can coexist: Ollama for everyday chat, with vLLM started on demand to process a batch of extractions. There is no reason to keep only one if the use cases differ. Just plan for the disk space needed to store the same model in two formats, such as a GGUF file for chat and Hugging Face weights for the server.
vLLM or llama.cpp: which should you choose?+
Does Ollama use llama.cpp?+
Can vLLM run GGUF files?+
Is ExLlamaV2 still maintained?+
Which engine is fastest?+
Do you need a GPU NVIDIA for vLLM?+
- Ollama vs. llama.cpp
- vLLM: the complete guide
- SGLang as a local LLM server
- MLX vs. llama.cpp on Mac
- llama.cpp: the complete guide
- Ollama, LM Studio, Jan, or GPT4All
- Source: llama.cpp repository
- Source: vLLM repository
- Source: ExLlamaV2 repository
- Source: Red Hat, Ollama versus vLLM
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.