Advanced 9 minLM Studio

MTP in LM Studio: enable Multi-Token Prediction

MTP (Multi-Token Prediction) is one of the few inference techniques that delivers a real tokens-per-second gain on LM Studio without degrading quality. Popularized by DeepSeek V3, it is now making its way into the llama.cpp runtime used by LM Studio. This guide explains what MTP is, which models it works on, how to enable it in LM Studio, and what gains to expect—without promising the moon.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
i
In brief
MTP (Multi-Token Prediction) generates multiple tokens per forward pass instead of one, increasing throughput without a faster card. · Only certain models support it natively: the DeepSeek line (V3, V3.1, V4 Flash) and some Qwen3 MoE variants with MTP heads in GGUF. · It is enabled in LM Studio 0.3.10+ through an up-to-date llama.cpp runtime, followed by a toggle in the model's Advanced Configuration. · Measured gain: DeepSeek V3 32B Q4 on RTX 4090 goes from about 25 to 42 tok/s on code.

#Why MTP speeds up inference

LLM inference is memory-bandwidth-bound: for every generated token, the model's entire set of weights must be reread from VRAM. On a 32B model in Q4, that's ~19 GB transferred at every step. Even a RTX 4090 tops out at around 1 TB/s of bandwidth, which mechanically caps throughput at a few dozen tokens per second.

MTP gets around this wall in a simple way: generate several tokens per forward pass instead of just one. If you can produce 2 or 3 valid tokens in a single memory round trip, throughput increases accordingly—without having to buy a faster card. That is what distinguishes MTP from other “optimizations” that merely make better use of compute (Flash Attention, prefix caching): here, it directly targets the dominant constraint.

i
The issue in one sentence
MTP doesn't make the model faster to load, smaller, or better. It simply does more useful work per VRAM access. On a local 32B model, that's the difference between 25 and 45 tok/s. That's huge.

#What MTP really does

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

In practice, a model trained with MTP has more than one prediction head. It has two, three, or four, in a cascade: the first predicts token N+1, the second attempts token N+2, and so on. During training, these auxiliary heads learn to anticipate subsequent tokens from intermediate hidden states. Once trained, the model can use these heads during inference to produce several candidates at once.

During generation, two patterns coexist in current runtimes:

MTP in speculative decoding
The auxiliary heads produce a draft of 2–4 tokens, then the main model verifies in a single forward pass whether to accept them. Any rejected token is regenerated. This is the default implementation in llama.cpp for DeepSeek V3.
MTP as an internal draft model
A variant where the secondary heads serve only as a small draft model, and the runtime treats it as a standard draft model, but hosted in the same GGUF file. Simpler to serve, but with a more modest gain.
→
Not magic, but solid
The acceptance rate of proposed tokens is the key factor. On structured text (code, JSON, markup), it exceeds 80% and the gain approaches 2×. On highly creative or unpredictable text, it sometimes falls below 50% and the gain drops to 1.2–1.3×.

#Which models actually support MTP

MTP is not a setting you can toggle on any model: the model must have been trained with MTP heads, and the corresponding weights must be present in the GGUF file. For now, the useful list for local use LM Studio is short:

DeepSeek V3 (base and chat)
The pioneer. 671B MoE architecture with 4 MTP heads. Locally, target aggressively quantized variants (Q2_K_XL, IQ3_XS) or distillations.
DeepSeek V3.1 / V3.2 / V4 Flash
The entire DeepSeek lineage inherited MTP. The Flash versions (≈300B MoE) are the only ones that truly fit on a Mac Studio Ultra or a multi-GPU build with 80–100 GB of VRAM.
Qwen3 MoE (variant selection)
Several Qwen3 MoE variants have released GGUF weights with MTP heads. Check the Hugging Face model page: the presence of an “mtp_layers” field or an -mtp suffix in the name is a good indicator.
Classic dense models (Gemma 4, Mistral Small, Qwen 3.5)
No native MTP. LM Studio won't do anything special. For these models, the only available equivalent is standard speculative decoding with a separate draft model.
!
Verify the GGUF
A GGUF file labeled “DeepSeek” does not necessarily contain the MTP heads: some community conversions removed them to save a few hundred MB. Look for an explicit “MTP” or “multi-token” mention in the repo description, or check the size—a model with MTP is 5-10% larger at the same quantization level.

#Requirements on the LM Studio side

recent LM Studio
Version 0.3.10 or later. Older versions use a llama.cpp runtime from before MTP support, and the option simply does not appear in the UI.
Up-to-date llama.cpp runtime
MTP support in llama.cpp arrived in late 2024 and was stabilized during 2025. LM Studio includes several runtime versions; you must explicitly select the latest one in the Runtimes tab.
Sufficient VRAM or unified memory
MTP adds little memory overhead (the auxiliary heads are small), but you still need to load the entire model. Allow for standard VRAM for the target size, plus 5–10% for the extended KV cache used during verification.
GPU detected correctly
The MTP gain appears only when inference is GPU-bound. In pure CPU mode, the benefit exists but is largely marginal and sometimes negative on small models.

#1. Enable the up-to-date runtime

Before touching the model, verify that LM Studio uses a version of llama.cpp that supports MTP. This is the most commonly overlooked step.

  1. 01
    Open Runtime settings
    In LM Studio, go to the gear icon (Settings), then to the “Runtimes” section (or “LM Runtimes,” depending on the version). You will see the installed runtimes there: CUDA, Metal, Vulkan, CPU.
  2. 02
    Install the latest version
    Click “Check for updates” next to your platform's runtime. LM Studio downloads a recent version of the bundled llama.cpp. MTP support is available in all 2025+ builds.
  3. 03
    Check the active version
    In the same section, the version currently in use is displayed (for example, “llama.cpp v0.3.x — CUDA”). Confirm that it is later than November 2024. If multiple runtimes coexist, select the newest one as the default.
i
The runtime, not the app
LM Studio and its llama.cpp runtime evolve independently. An old version of LM Studio can work perfectly well with a recent runtime: what matters is the version of the runtime bundled for MTP.

#2. Load a model with an MTP head

Download an MTP-compatible model through the built-in search tab. Type “DeepSeek V3” or “Qwen3 MoE” and filter for recent GGUF builds that mention MTP. If necessary, go directly through Hugging Face and drag the GGUF file into the LM Studio model folder.

Model folder path LM Studio
# Linux / macOS
~/.lmstudio/models/

# Windows
%USERPROFILE%\.lmstudio\models\

Once the model appears in the library, load it normally from the Chat or Local Server tab. Before clicking “Load,” open the “Advanced Configuration” panel: that is where the MTP options appear when the model declares them.

→
The good sign
If the MTP option does not appear in Advanced Configuration, it is almost always because the MTP heads were not included in the GGUF. Download another variant before wasting time looking for the problem.

#3. Tune the MTP parameters

Advanced Configuration generally exposes three MTP-related parameters. The names vary depending on the LM Studio version, but the principle remains the same.

Enable MTP (or Use MTP heads)
Main switch. Enable it. When disabled, the model ignores its auxiliary heads and behaves like a standard model.
MTP draft tokens (n_draft)
Number of tokens proposed per pass. Typical value: 3 or 4. Beyond that, the acceptance rate drops and the gain evaporates. Below that (1-2), the verification cost becomes comparable to the gain.
MTP acceptance threshold
Confidence threshold for accepting a proposed token. Often around 0.7 by default. Higher = more safety but less gain. Lower = faster, but output quality may subtly drift.
Example of a saved config (model_config.json)
{
  "mtp": {
    "enabled": true,
    "n_draft": 3,
    "acceptance_threshold": 0.7
  },
  "gpu_layers": -1,
  "context_length": 8192
}
!
Don't lower the threshold without testing
Lowering acceptance_threshold to 0.5 gains 10-15% tokens/sec but introduces minor style drift in long generations. Compare your outputs side by side before locking in this setting.

#4. Measure the real-world gain on your machine

Enabling MTP without measuring means missing half the benefit described in the guide. LM Studio displays the throughput at the bottom of the chat (tok/s) after each generation. The honest method:

  1. 01
    Prepare 3 representative prompts
    A code prompt (highly structured, with a high acceptance rate expected), a factual French prompt, and an open-ended creative prompt. Save them so you can replay them identically.
  2. 02
    Disable MTP, measure the baseline
    Unload the model, disable MTP, and reload it. Run each prompt and record the average tok/s over 3 repetitions. Do not forget to ignore the very first run (the prefill is not comparable).
  3. 03
    Enable MTP, measure
    Unload it, reactivate MTP, reload it with exactly the same parameters (context, temperature, etc.). Repeat the same prompts.
  4. 04
    Compare
    Calculate the ratio per prompt. A factor of 1.5-2× is expected for code, 1.3-1.7× for factual French, and 1.1-1.4× for creative work. Below 1.2× across the board, something is wrong (runtime too old, missing MTP heads, or GPU saturated by something else).
DeepSeek distilled V3 32B Q4, RTX 4090
Baseline ~25 tok/s, MTP enabled ~42 tok/s on a Python code prompt. On French prose, ~30 tok/s.
DeepSeek V4 Flash on Mac Studio M2 Ultra 192 GB
Baseline ~14 tok/s, ~24 tok/s with MTP enabled on code. On free-form text, ~18 tok/s. The gain is significant even on Apple Silicon.
Qwen3 MoE 30B-A3B Q4, RTX 4070 Ti
Baseline ~38 tok/s, MTP enabled ~55 tok/s on JSON. On creative text, the gain is limited to ~46 tok/s.

#Advanced settings and combinations

MTP combines well with other optimizations available in LM Studio. Some combinations that work:

MTP + Flash Attention
Enable them at the same time. Flash Attention reduces the KV-cache memory cost during draft verification, which helps when n_draft is high.
MTP + Q4_K_M (vs Q8)
MTP preserves quality even with aggressive Q4. There’s no reason to move up to Q8 “to compensate” — keep Q4_K_M and save VRAM.
MTP + long context
The KV cache is heavier with MTP (verification stores more intermediate states). For a 32B model with a 32k context, expect an additional 1–2 GB of VRAM.
MTP + batch (Local Server)
If you serve multiple simultaneous requests through LM Studio's OpenAI-compatible server, the MTP gain compounds with the batching gain. That's where you see the biggest differences.
→
Profile by use case
For code assistance (Continue.dev, Aider) connected to LM Studio, increase n_draft to 4: code is so predictable that the acceptance rate remains high. For general-purpose chat, stay at 3. For fiction or brainstorming, drop to 2 or disable MTP.

#Troubleshooting

The MTP option doesn’t appear
Either the runtime is too old (check the version under Settings → Runtimes), or the GGUF model doesn't contain MTP heads. Download a variant explicitly marked MTP.
MTP enabled, but zero measured gain
Make sure the GPU is actually the target (gpu_layers = -1 or equal to the total number of layers). On pure CPU, MTP sometimes costs more than it delivers. Also check that no other process is saturating VRAM.
Drifting output or strange repetitions
Acceptance threshold too low. Raise it to 0.75-0.8. If the problem persists, disable MTP: pathological cases exist with highly atypical prompts.
LM Studio crashes while loading
A poorly converted GGUF with corrupted MTP heads can crash the runtime. Try another quantization of the same model, or a file from a different publisher on Hugging Face.
Noticeable improvement in chat but not through the API
Make sure the Local Server is using the same configuration as the chat—there is a separate settings page for the server, where MTP must be re-enabled independently.
i
When to keep MTP disabled
For custom fine-tunes where you suspect the auxiliary heads did not learn properly, on a small CPU-only model, or while debugging when you want the most deterministic output possible: disabling MTP is a reasonable decision.

#Go further

MTP is one of the performance levers that really works locally. A few additional avenues to push things further:

Get started with LM Studio
If you arrived here without installing LM Studio, the getting-started guide covers the basics before you touch the advanced settings.
Turn LM Studio into an API server
To expose an OpenAI-compatible endpoint and benefit from MTP across all your clients (Continue.dev, Aider, Python scripts).
LM Studio plugins: which ones to install and how
Beyond MTP, the plugin ecosystem and the lmstudio-js/python SDK provide additional flexibility.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.