MTP in LM Studio: enable Multi-Token Prediction
MTP (Multi-Token Prediction) is one of the few inference techniques that delivers a real tokens-per-second gain on LM Studio without degrading quality. Popularized by DeepSeek V3, it is now making its way into the llama.cpp runtime used by LM Studio. This guide explains what MTP is, which models it works on, how to enable it in LM Studio, and what gains to expect—without promising the moon.
#Why MTP speeds up inference
LLM inference is memory-bandwidth-bound: for every generated token, the model's entire set of weights must be reread from VRAM. On a 32B model in Q4, that's ~19 GB transferred at every step. Even a RTX 4090 tops out at around 1 TB/s of bandwidth, which mechanically caps throughput at a few dozen tokens per second.
MTP gets around this wall in a simple way: generate several tokens per forward pass instead of just one. If you can produce 2 or 3 valid tokens in a single memory round trip, throughput increases accordingly—without having to buy a faster card. That is what distinguishes MTP from other “optimizations” that merely make better use of compute (Flash Attention, prefix caching): here, it directly targets the dominant constraint.
#What MTP really does
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
In practice, a model trained with MTP has more than one prediction head. It has two, three, or four, in a cascade: the first predicts token N+1, the second attempts token N+2, and so on. During training, these auxiliary heads learn to anticipate subsequent tokens from intermediate hidden states. Once trained, the model can use these heads during inference to produce several candidates at once.
During generation, two patterns coexist in current runtimes:
- MTP in speculative decoding
- The auxiliary heads produce a draft of 2–4 tokens, then the main model verifies in a single forward pass whether to accept them. Any rejected token is regenerated. This is the default implementation in llama.cpp for DeepSeek V3.
- MTP as an internal draft model
- A variant where the secondary heads serve only as a small draft model, and the runtime treats it as a standard draft model, but hosted in the same GGUF file. Simpler to serve, but with a more modest gain.
#Which models actually support MTP
MTP is not a setting you can toggle on any model: the model must have been trained with MTP heads, and the corresponding weights must be present in the GGUF file. For now, the useful list for local use LM Studio is short:
- DeepSeek V3 (base and chat)
- The pioneer. 671B MoE architecture with 4 MTP heads. Locally, target aggressively quantized variants (Q2_K_XL, IQ3_XS) or distillations.
- DeepSeek V3.1 / V3.2 / V4 Flash
- The entire DeepSeek lineage inherited MTP. The Flash versions (≈300B MoE) are the only ones that truly fit on a Mac Studio Ultra or a multi-GPU build with 80–100 GB of VRAM.
- Qwen3 MoE (variant selection)
- Several Qwen3 MoE variants have released GGUF weights with MTP heads. Check the Hugging Face model page: the presence of an “mtp_layers” field or an -mtp suffix in the name is a good indicator.
- Classic dense models (Gemma 4, Mistral Small, Qwen 3.5)
- No native MTP. LM Studio won't do anything special. For these models, the only available equivalent is standard speculative decoding with a separate draft model.
#Requirements on the LM Studio side
- recent LM Studio
- Version 0.3.10 or later. Older versions use a llama.cpp runtime from before MTP support, and the option simply does not appear in the UI.
- Up-to-date llama.cpp runtime
- MTP support in llama.cpp arrived in late 2024 and was stabilized during 2025. LM Studio includes several runtime versions; you must explicitly select the latest one in the Runtimes tab.
- Sufficient VRAM or unified memory
- MTP adds little memory overhead (the auxiliary heads are small), but you still need to load the entire model. Allow for standard VRAM for the target size, plus 5–10% for the extended KV cache used during verification.
- GPU detected correctly
- The MTP gain appears only when inference is GPU-bound. In pure CPU mode, the benefit exists but is largely marginal and sometimes negative on small models.
#1. Enable the up-to-date runtime
Before touching the model, verify that LM Studio uses a version of llama.cpp that supports MTP. This is the most commonly overlooked step.
- 01Open Runtime settingsIn LM Studio, go to the gear icon (Settings), then to the “Runtimes” section (or “LM Runtimes,” depending on the version). You will see the installed runtimes there: CUDA, Metal, Vulkan, CPU.
- 02Install the latest versionClick “Check for updates” next to your platform's runtime. LM Studio downloads a recent version of the bundled llama.cpp. MTP support is available in all 2025+ builds.
- 03Check the active versionIn the same section, the version currently in use is displayed (for example, “llama.cpp v0.3.x — CUDA”). Confirm that it is later than November 2024. If multiple runtimes coexist, select the newest one as the default.
#2. Load a model with an MTP head
Download an MTP-compatible model through the built-in search tab. Type “DeepSeek V3” or “Qwen3 MoE” and filter for recent GGUF builds that mention MTP. If necessary, go directly through Hugging Face and drag the GGUF file into the LM Studio model folder.
Once the model appears in the library, load it normally from the Chat or Local Server tab. Before clicking “Load,” open the “Advanced Configuration” panel: that is where the MTP options appear when the model declares them.
#3. Tune the MTP parameters
Advanced Configuration generally exposes three MTP-related parameters. The names vary depending on the LM Studio version, but the principle remains the same.
- Enable MTP (or Use MTP heads)
- Main switch. Enable it. When disabled, the model ignores its auxiliary heads and behaves like a standard model.
- MTP draft tokens (n_draft)
- Number of tokens proposed per pass. Typical value: 3 or 4. Beyond that, the acceptance rate drops and the gain evaporates. Below that (1-2), the verification cost becomes comparable to the gain.
- MTP acceptance threshold
- Confidence threshold for accepting a proposed token. Often around 0.7 by default. Higher = more safety but less gain. Lower = faster, but output quality may subtly drift.
#4. Measure the real-world gain on your machine
Enabling MTP without measuring means missing half the benefit described in the guide. LM Studio displays the throughput at the bottom of the chat (tok/s) after each generation. The honest method:
- 01Prepare 3 representative promptsA code prompt (highly structured, with a high acceptance rate expected), a factual French prompt, and an open-ended creative prompt. Save them so you can replay them identically.
- 02Disable MTP, measure the baselineUnload the model, disable MTP, and reload it. Run each prompt and record the average tok/s over 3 repetitions. Do not forget to ignore the very first run (the prefill is not comparable).
- 03Enable MTP, measureUnload it, reactivate MTP, reload it with exactly the same parameters (context, temperature, etc.). Repeat the same prompts.
- 04CompareCalculate the ratio per prompt. A factor of 1.5-2× is expected for code, 1.3-1.7× for factual French, and 1.1-1.4× for creative work. Below 1.2× across the board, something is wrong (runtime too old, missing MTP heads, or GPU saturated by something else).
- DeepSeek distilled V3 32B Q4, RTX 4090
- Baseline ~25 tok/s, MTP enabled ~42 tok/s on a Python code prompt. On French prose, ~30 tok/s.
- DeepSeek V4 Flash on Mac Studio M2 Ultra 192 GB
- Baseline ~14 tok/s, ~24 tok/s with MTP enabled on code. On free-form text, ~18 tok/s. The gain is significant even on Apple Silicon.
- Qwen3 MoE 30B-A3B Q4, RTX 4070 Ti
- Baseline ~38 tok/s, MTP enabled ~55 tok/s on JSON. On creative text, the gain is limited to ~46 tok/s.
#Advanced settings and combinations
MTP combines well with other optimizations available in LM Studio. Some combinations that work:
- MTP + Flash Attention
- Enable them at the same time. Flash Attention reduces the KV-cache memory cost during draft verification, which helps when n_draft is high.
- MTP + Q4_K_M (vs Q8)
- MTP preserves quality even with aggressive Q4. There’s no reason to move up to Q8 “to compensate” — keep Q4_K_M and save VRAM.
- MTP + long context
- The KV cache is heavier with MTP (verification stores more intermediate states). For a 32B model with a 32k context, expect an additional 1–2 GB of VRAM.
- MTP + batch (Local Server)
- If you serve multiple simultaneous requests through LM Studio's OpenAI-compatible server, the MTP gain compounds with the batching gain. That's where you see the biggest differences.
#Troubleshooting
- The MTP option doesn’t appear
- Either the runtime is too old (check the version under Settings → Runtimes), or the GGUF model doesn't contain MTP heads. Download a variant explicitly marked MTP.
- MTP enabled, but zero measured gain
- Make sure the GPU is actually the target (gpu_layers = -1 or equal to the total number of layers). On pure CPU, MTP sometimes costs more than it delivers. Also check that no other process is saturating VRAM.
- Drifting output or strange repetitions
- Acceptance threshold too low. Raise it to 0.75-0.8. If the problem persists, disable MTP: pathological cases exist with highly atypical prompts.
- LM Studio crashes while loading
- A poorly converted GGUF with corrupted MTP heads can crash the runtime. Try another quantization of the same model, or a file from a different publisher on Hugging Face.
- Noticeable improvement in chat but not through the API
- Make sure the Local Server is using the same configuration as the chat—there is a separate settings page for the server, where MTP must be re-enabled independently.
#Go further
MTP is one of the performance levers that really works locally. A few additional avenues to push things further:
- Get started with LM Studio
- If you arrived here without installing LM Studio, the getting-started guide covers the basics before you touch the advanced settings.
- Turn LM Studio into an API server
- To expose an OpenAI-compatible endpoint and benefit from MTP across all your clients (Continue.dev, Aider, Python scripts).
- LM Studio plugins: which ones to install and how
- Beyond MTP, the plugin ecosystem and the lmstudio-js/python SDK provide additional flexibility.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.