TurboQuant: quantize large models to fit them in local
TurboQuant is a recent quantization method that pushes the size/quality tradeoff inherited from GGUF K-quants further. The goal: fit a dense 70B or a 100B+ MoE on a RTX 4090 24 GB without quality collapsing. This guide explains the principle, measures the real gains, and provides the concrete workflow for quantizing a Hugging Face model yourself.
#Why TurboQuant?
GGUF K-quants (Q4_K_M, Q5_K_M…) have been the standard since 2023: universally supported by llama.cpp, Ollama, LM Studio, and easy to produce. But they reach their limit on very large models. A dense 70B model in Q4_K_M requires about 40 GB of VRAM—out of reach for a RTX 4090 24 GB. Dropping to Q3 or Q2 causes quality to collapse.
Turboquant quantization specifically addresses this gap. It is a family of techniques that combine dataset calibration, non-uniform bit allocation by layer, and block compression with a shared dictionary. The result: 2.5 to 3 effective bits per weight, with less quality loss than a Q3_K_M GGUF.
#How it compresses
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Three mechanisms add up. None is revolutionary on its own; it’s their combination that makes the difference.
- Calibration on data
- A few hundred representative prompts are run through the model to measure which weights actually influence the outputs. The others can be quantized aggressively.
- Bit allocation by layer
- Instead of using a fixed budget (4 bits everywhere), TurboQuant assigns 5–6 bits to critical attention layers and drops to 2 bits for redundant MLPs. The average total falls to 2.7–3.2 bpw.
- Block compression
- The weights are grouped into blocks of 32 or 64 values that share a scale and offset, with a codebook compressed per group. This avoids the per-parameter overhead of K-quants.
#TurboQuant vs. classic GGUF
For a dense 70B model, here is the typical order of magnitude:
- GGUF Q4_K_M (~4.5 bpw)
- ~40 GB, near-FP16 quality, supported everywhere. The reference standard.
- GGUF Q3_K_M (~3.4 bpw)
- ~32 GB, noticeable loss in long-form reasoning, with hallucinations appearing.
- GGUF Q2_K (~2.6 bpw)
- ~26 GB, clearly degraded model, avoid for serious use.
- TurboQuant ~3.0 bpw
- ~28 GB, with quality close to Q4_K_M. Fits on 1×RTX 3090 or 1×4090 with a short context.
- TurboQuant ~2.5 bpw
- ~23 GB, slightly below Q4 but well above Q3. Enables 70B on 24 GB of VRAM.
#Measured VRAM savings by model size
The figures below are rough estimates for recent dense models (Qwen 3.5, Granite 4.2), with a 4k context and no Flash Attention. Add 10-20% headroom for the KV cache and runtime overhead.
- 7B — Q4_K_M: 4.4 GB
- TurboQuant 3.0 bpw: ~2.9 GB. Limited benefit; the 7B already fits everywhere.
- 13B — Q4_K_M: 7.8 GB
- TurboQuant 3.0 bpw: ~5.1 GB. Enables 16k+ context on 8 GB of VRAM.
- 32B — Q4_K_M: 19 GB
- TurboQuant 3.0 bpw: ~13 GB. Fits comfortably on RTX 4080 16 GB with an 8k context.
- 70B — Q4_K_M: 40 GB
- TurboQuant 2.5 bpw: ~23 GB. Enables 70B on 1×RTX 4090 24 GB or 1×3090 24 GB.
- 120B MoE — Q4_K_M: ~70 GB
- TurboQuant 2.7 bpw: ~42 GB. Feasible on 2×3090 or a 64 GB Mac Studio.
#Impact on quality
On the usual benchmarks (MMLU, HellaSwag, HumanEval), a 3.0 bpw TurboQuant version of a 70B remains within 0.5–1.5 points of FP16. On multi-step reasoning and long-form code, clearer differences begin to appear. The test that separates the methods most clearly: a long Python code excerpt with cross-dependencies.
- General knowledge
- Almost imperceptible. If you ask general-knowledge questions, you won't notice the difference.
- Short reasoning
- Minor loss. The chains of thought remain coherent over 5–10 steps.
- Long reasoning (math/proofs)
- Noticeable degradation. The model may skip a step or lose track after 20+ reasoning turns. Favor 3.5+ bpw if that is your use case.
- Code
- Sensitive to aggressive quantization. Below 3.0 bpw, expect more subtle bugs (wrong indexes, reversed conditions). Stick with Q4_K_M or TurboQuant ≥3.5 bpw for Aider/Continue.
- Rare languages
- French works well down to 2.5 bpw. Languages that are underrepresented in the calibration corpus suffer more.
#Hardware and software requirements
Quantizing a model yourself is still a demanding task. Downloading a pre-quantized version from Hugging Face is almost always simpler. If you still want to produce your own version:
- GPU for calibration
- You must be able to load the model in FP16 or BF16 during analysis. For a 70B, that means 2×A100 80 GB or a 192 GB Mac Studio M2 Ultra. For a 32B, a RTX 4090 24 GB is sufficient for partial offload.
- Base model
- The unquantized weights (safetensors) downloaded from the original Hugging Face repo. Allow 140 GB for a 70B FP16 model.
- Calibration dataset
- 256 to 1024 representative samples of your use case. French Wikipedia, code, and your own prompts. Avoid generic datasets if you have a specific domain.
- Target runtime
- Decide in advance: ExLlamaV3 (fastest on Nvidia), llama.cpp with the turbo backend, or a specific fork. The file format differs between runtimes.
- Python 3.10+ and CUDA 12+
- Standard toolchain. On AMD, ROCm 6.x works with llama.cpp but not yet with all turbo forks.
#Workflow for quantizing it yourself
The typical process, from an FP16 Hugging Face model to an executable file in Ollama or llama.cpp.
- 01Retrieve the original weightsClone the Hugging Face repo for the unquantized model. Allow at least an hour for a 70B model with a good connection. Use huggingface-cli download so you can resume an interrupted download.
- 02Preparing the calibration datasetBuild a JSONL file with 256–1024 prompts. Diversity matters more than quantity: code, French prose, dialogues, technical questions. A 5–20 MB file is more than enough.
- 03Start calibrationThe tool (for example, the quantize script provided by the selected TurboQuant project) loads the model, runs each prompt forward, and collects layer-by-layer activation statistics. Allow 1-4h on a 70B, depending on the GPU.
- 04Quantify and exportBit allocation is calculated from the statistics, then each tensor is encoded. Output: a .safetensors or .gguf file, depending on the runtime. Allow an additional 30 minutes to 2 hours.
- 05Convert to the runtime formatFor Ollama: create a Modelfile pointing to the quantized file, then ollama create monmodele -f Modelfile. For llama.cpp: the main binary accepts the .gguf directly.
- 06Validate the qualityRun a small test suite: 20–30 prompts covering your use cases. Compare side by side with the reference GGUF Q4_K_M version. If the degradation is too severe, repeat with more bits or a better dataset.
#Common pitfalls and troubleshooting
- The runtime does not recognize the file
- TurboQuant is not a single standard format. Make sure you're loading the file with the runtime that corresponds to the method used (ExLlamaV3 for EXL3, a recent version of llama.cpp for turbo GGUFs).
- VRAM exceeded during inference even though the file fit
- You’re forgetting the KV cache. With a 32k context, the KV cache can take up an additional 4–8 GB. Reduce num_ctx or enable Flash Attention with OLLAMA_FLASH_ATTENTION=1.
- Catastrophic quality for your use case
- The calibration dataset does not cover your domain. Re-run calibration with 100–200 representative prompts. This is by far the most powerful lever.
- The quantized model is slower than Q4_K_M
- Normal on some architectures: turbo decompression has overhead. On older cards (RTX 20xx), the VRAM gain may come at the cost of tokens/sec. Measure before adopting it.
- Differences between runs
- Calibration introduces nondeterminism. Two quantizations of the same model with the same dataset may differ slightly in quality. Run several times and keep the best one.
#Go further
TurboQuant is one of several tools for pushing the limit of what fits locally. A few complementary options:
- Choose your quantization (Q4, Q5, Q8, FP16)
- To properly position TurboQuant against classic GGUF K-quants and choose on a case-by-case basis.
- Local LLM fine-tuning: LoRA and QLoRA
- If you want to go beyond quantization and adapt a model to your domain—often a better lever than more aggressive quantization.
- Choose your GPU for local AI
- To decide which card to target, knowing that VRAM remains the number-one constraint, even with TurboQuant.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.