Advanced 11 minBackends

llama.cpp vs vLLM vs Exllama

Direct response

llama.cpp is the portable engine for personal use (GGUF, CPU, Mac, GPUs of all brands) and for Ollama and LM Studio. vLLM is designed to serve many users simultaneously on GPUs, with continuous batching; Red Hat measures up to 793 tokens/s versus 41 for Ollama on an A100. ExLlamaV2 is archived: its development continues in ExLlamaV3.

Under Ollama, LM Studio, or Jan is an inference engine, and that engine determines the accepted formats, usable hardware, and behavior under load. This guide compares llama.cpp, vLLM, ExLlama, and SGLang using criteria that can be verified in their repositories, corrects misconceptions, and provides a selection rule. It publishes no in-house throughput figures: only attributed third-party measurements appear here.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#An inference engine: what really changes between them

An inference engine transforms input tokens into output tokens. The applications you install (Ollama, LM Studio, Jan) include one; production servers (vLLM, SGLang) are one. Three differences affect how you use them: the model formats they accept, the hardware they can use, and how they handle multiple requests at the same time.

The engines in one table (official repositories, September 2026)
EngineLicenseAnnounced formats and hardwareTypical use
llama.cppMITGGUF; CPU, Apple Silicon, NVIDIA (CUDA), AMD (HIP), Vulkan, SYCL, and othersPersonal workstation, laptop, lightweight server
vLLMApache 2.0Hugging Face models; FP8, INT4, GPTQ, AWQ, GGUF (experimental); NVIDIA GPU, AMD, Intel, and x86/ARM CPUsServe many users, production
SGLangApache 2.0Hugging Face models; FP4, FP8, INT4, AWQ, GPTQHigh-throughput server, shared prefixes
ExLlamaV3MITEXL3; consumer GPU NVIDIALatency on a personal GPU with TabbyAPI
MLX LMMITApple Silicon onlyMac: generation and fine-tuning
i
Ollama and LM Studio are not competing engines to these
The Ollama repository lists llama.cpp under “Supported backends,” and LM Studio’s homepage indicates an engine based on MLX and llama.cpp. Comparing “Ollama vs llama.cpp” is like comparing an application with its engine: the dedicated guide Ollama vs llama.cpp covers this case.

#The model format determines the possible engine

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before comparing speeds, check the format in which the model you want is published. A GGUF file loads in llama.cpp and therefore in Ollama, LM Studio, and Jan. Hugging Face weights in FP16 or FP8, or quantized in AWQ or GPTQ, load in vLLM or SGLang. A model converted for MLX loads with MLX LM on Mac, and an EXL3 model with ExLlamaV3. Switching engines often means downloading the model again in another format.

Format and compatible engine
FormatNatural engineAnother possible engine
GGUFllama.cpp (and therefore Ollama, LM Studio, Jan)vLLM, in experimental mode
Hugging Face weights (FP16, FP8)vLLM, SGLangMLX LM after conversion, on Mac
AWQ, GPTQvLLM, SGLangCheck depending on the engine
EXL3ExLlamaV3None
MLXMLX LMLM Studio and Jan, which announce MLX

#llama.cpp: the portable standard

llama.cpp is a dependency-free C/C++ implementation whose stated goal is inference with minimal configuration across a wide range of hardware. Apple Silicon is prioritized (NEON, Accelerate, Metal), x86 CPUs are supported with AVX, AVX2, AVX512, and AMX, and NVIDIA (CUDA), AMD (HIP), Vulkan, and SYCL GPUs are supported. It offers 1.5- to 8-bit quantization and hybrid CPU/GPU inference, allowing you to run a model larger than the available VRAM.

A misconception to correct: llama.cpp is not limited to one request at a time. Its server advertises multi-user parallel generation and continuous batching, enabled by default, with slots configured by the --parallel parameter. It therefore serves a small team effectively; what it does not aim to match are vLLM’s large-scale production features.

Strengths
Portability (Mac, Linux, Windows, Android), ubiquitous GGUF formats, CPU and GPU hybrid operation, low dependency.
Limitations
No distributed deployment features comparable to those in vLLM; settings (context, slots, layers on the GPU) to know.
Ecosystem
Foundation for Ollama and LM Studio, also cited by Jan.

For compilation with CUDA, Metal, or Vulkan, see the compilation guides; the complete guide to llama.cpp details how to use it.

#vLLM: the server for many users

vLLM is an inference and serving library created at the Sky Computing Lab at the University of California, Berkeley. It is built on PagedAttention, which manages attention key and value memory in pages, and continuous batching, complemented by chunked prefill and prefix caching. It advertises FP8, INT4, GPTQ, AWQ, and GGUF formats, an OpenAI-compatible API, and support for NVIDIA, AMD, and Intel GPUs, as well as x86, ARM, and PowerPC CPUs.

Two misconceptions to correct. First, vLLM is no longer limited to NVIDIA or CPU-less setups: its repository indicates support for multiple GPUs and CPUs. Second, it is no longer without GGUF support: its documentation describes that support as highly experimental and poorly optimized, useful mainly for reducing memory usage. For everyday GGUF use, llama.cpp remains the standard approach.

Strengths
Throughput under load, page-managed cache memory, OpenAI-compatible API, parallelism (tensor, pipeline, expert), and numerous formats.
Limitations
Heavier to install and configure than llama.cpp; designed for GPU servers; GGUF is not its domain.
Ecosystem
The default choice when multiple people or applications query the same model at the same time.

The guide on vLLM explains what the engine is; the production deployment guide covers configuration and monitoring.

#ExLlama: V2 archived, V3 in development

The ExLlamaV2 repository notes that the project is currently archived and development continues in ExLlamaV3. Many comparisons, including the previous version of this page, still present V2 as the state-of-the-art option. The ExLlamaV3 repository announces the EXL3 quantization format, tensor and expert parallel inference, CPU offloading for mixture-of-experts models, continuous batching, speculative decoding, and an OpenAI-compatible API through TabbyAPI, its recommended server.

ExLlamaV3 targets consumer GPUs, not production servers or Macs. If you have a NVIDIA card and want the best balance between quality and size, EXL3 quantization is worth trying; first check that your model appears in the repository’s list of supported architectures.

#SGLang, MLX LM, and others

SGLang
A serving framework described as targeting low latency and high throughput, from one GPU to large clusters. It advertises RadixAttention for prefix caching, continuous batching, PagedAttention, and speculative decoding. It competes with vLLM on servers; the dedicated guide covers it in detail.
MLX LM
Python package for generating text and fine-tuning models on Apple Silicon with MLX. It doesn't work outside of Mac. The MLX versus llama.cpp guide compares the two on Mac.
TensorRT-LLM
NVIDIA engine, highly performant on its cards but more demanding to set up; consider it only for a production NVIDIA fleet.

#What the published measurements say

Throughput depends on the hardware, model, quantization, query length, and engine version, and changes every month. This guide therefore presents no in-house measurements and advises you to be wary of tokens-per-second tables without a stated protocol. One third-party source does publish a complete protocol: Red Hat, in August 2025.

Red Hat measurement: vLLM versus Ollama on an A100 (August 2025)
ItemValue published by Red Hat
HardwareOne NVIDIA A100-PCIE-40GB card
ModelLlama 3.1 8B Instruct (FP16 on the Ollama side)
VersionsvLLM 0.9.1; Ollama 0.9.2
Test toolGuideLLM 0.2.1, from 1 to 256 concurrent users
Maximum throughput793 tokens/s for vLLM versus 41 for Ollama
Peak P99 latency80 ms for vLLM versus 673 ms for Ollama

Read with reservations: the article is associated with Red Hat AI products, so it comes from a commercial player in the field; the tests cover Ollama rather than llama.cpp directly, using default settings; and they date back several versions. The robust takeaway is qualitative: under heavy concurrency, a serving engine with aggressive batching outperforms an application designed for a single user. For solo use, throughput is not what will hold you back.

!
No tokens-per-second table on this page
The previous version of this guide provided throughput and time-to-first-response figures for three engines. They were not tied to any source and have been removed. To compare on your machine, measure them yourself with your model and queries.

#Memory: what each engine reserves

A model’s memory consists of its weights, context cache (KV), and headroom. The weights are easy to calculate: an 8-billion-parameter model uses approximately 16 GB in FP16 (8 billion times 2 bytes) and approximately 5 GB in Q4, according to the site’s rule of thumb. The context cache grows with conversation length and the number of simultaneous requests.

How each engine handles context memory
EngineDocumented behaviorPractical consequence
llama.cppUser-set context; parallel slots with a unified cacheYou set the context size and number of slots
vLLMPre-allocates part of GPU memory for the cache, 92% by defaultOn a 24 GB card, approximately 22 GB are reserved immediately
ExLlamaV32- to 8-bit cache quantizationThe cache can be compressed to fit in VRAM

The key point: vLLM reserves memory from the start, making it effective for serving but poorly suited to a card shared with other applications. Reduce gpu_memory_utilization if the card is also used for something else.

#Which engine for which use case

Decision by situation
Your situationRecommended engineReason
Personal use on Mac, PC, or laptopllama.cpp, via Ollama or LM StudioPortable and simple
Mac Apple Silicon, throughput-focusedMLX LM or llama.cppMLX is designed for Apple Silicon; compare on your models
A team or application queries the modelvLLM or SGLangContinuous batching, cache pages
A consumer-grade NVIDIA card, maximum qualityExLlamaV3 with TabbyAPIEXL3 quantization
Mixed hardware, CPU, AMD, Intelllama.cppBroad hardware support
GGUF-only modelllama.cppvLLM supports it only experimentally

The engines can coexist: Ollama for everyday chat, with vLLM started on demand to process a batch of extractions. There is no reason to keep only one if the use cases differ. Just plan for the disk space needed to store the same model in two formats, such as a GGUF file for chat and Hugging Face weights for the server.

Frequently asked questions about inference engines
vLLM or llama.cpp: which should you choose?+
llama.cpp for personal use or varied hardware, including CPUs and Macs. vLLM for serving multiple users on GPUs, thanks to continuous batching and paged memory management. If you are the only person using your machine, llama.cpp, through Ollama or LM Studio, is almost always sufficient.
Does Ollama use llama.cpp?+
The Ollama repository lists llama.cpp under “Supported backends.” Ollama adds model management, an API, and an application on top. Comparing Ollama and llama.cpp therefore means comparing an application with its engine, with different defaults, formats, and features; see the dedicated guide to this comparison for details.
Can vLLM run GGUF files?+
Yes, but its documentation describes this support as highly experimental and poorly optimized, useful mainly for reducing memory usage. It now goes through a separate plugin. For vLLM, prefer Hugging Face weights in FP16, FP8, or AWQ or GPTQ quantization, and keep GGUF for llama.cpp.
Is ExLlamaV2 still maintained?+
No: its repository currently shows that it is archived and that development continues in ExLlamaV3. ExLlamaV3 provides the EXL3 format and a recommended server, TabbyAPI. If you started from a tutorial about ExLlamaV2 and the EXL2 format, look for the V3 equivalent before you begin.
Which engine is fastest?+
It depends on the hardware, model, and number of simultaneous users. Under heavy concurrency, Red Hat measures a clear advantage for vLLM over Ollama. For one request at a time, the differences are smaller and depend on the hardware. Measure with your own model before deciding.
Do you need a GPU NVIDIA for vLLM?+
No. The vLLM repository announces support for NVIDIA, AMD, and Intel GPUs, as well as x86, ARM, and PowerPC CPUs, with extensions for other accelerators. Feature coverage varies by hardware: consult your platform's installation documentation before committing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.