Advanced 12 minQuantization

TurboQuant: quantize large models to fit them in local

TurboQuant is a recent quantization method that pushes the size/quality tradeoff inherited from GGUF K-quants further. The goal: fit a dense 70B or a 100B+ MoE on a RTX 4090 24 GB without quality collapsing. This guide explains the principle, measures the real gains, and provides the concrete workflow for quantizing a Hugging Face model yourself.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why TurboQuant?

GGUF K-quants (Q4_K_M, Q5_K_M…) have been the standard since 2023: universally supported by llama.cpp, Ollama, LM Studio, and easy to produce. But they reach their limit on very large models. A dense 70B model in Q4_K_M requires about 40 GB of VRAM—out of reach for a RTX 4090 24 GB. Dropping to Q3 or Q2 causes quality to collapse.

Turboquant quantization specifically addresses this gap. It is a family of techniques that combine dataset calibration, non-uniform bit allocation by layer, and block compression with a shared dictionary. The result: 2.5 to 3 effective bits per weight, with less quality loss than a Q3_K_M GGUF.

i
A family, not a single format
The term “TurboQuant” covers several implementations (AWQ-turbo, EXL3, HQQ+, and their variants). They all share the same intuition: measure the importance of weights by activation and allocate bits accordingly. The files are not interchangeable between runtimes.

#How it compresses

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Three mechanisms add up. None is revolutionary on its own; it’s their combination that makes the difference.

Calibration on data
A few hundred representative prompts are run through the model to measure which weights actually influence the outputs. The others can be quantized aggressively.
Bit allocation by layer
Instead of using a fixed budget (4 bits everywhere), TurboQuant assigns 5–6 bits to critical attention layers and drops to 2 bits for redundant MLPs. The average total falls to 2.7–3.2 bpw.
Block compression
The weights are grouped into blocks of 32 or 64 values that share a scale and offset, with a codebook compressed per group. This avoids the per-parameter overhead of K-quants.
→
Why it works
LLMs are heavily overparameterized: roughly 10-15% of the weights carry most of the signal. Identifying and preserving this subset lets you aggressively prune the rest without breaking the model.

#TurboQuant vs. classic GGUF

For a dense 70B model, here is the typical order of magnitude:

GGUF Q4_K_M (~4.5 bpw)
~40 GB, near-FP16 quality, supported everywhere. The reference standard.
GGUF Q3_K_M (~3.4 bpw)
~32 GB, noticeable loss in long-form reasoning, with hallucinations appearing.
GGUF Q2_K (~2.6 bpw)
~26 GB, clearly degraded model, avoid for serious use.
TurboQuant ~3.0 bpw
~28 GB, with quality close to Q4_K_M. Fits on 1×RTX 3090 or 1×4090 with a short context.
TurboQuant ~2.5 bpw
~23 GB, slightly below Q4 but well above Q3. Enables 70B on 24 GB of VRAM.
i
Reading the table
At the same size, TurboQuant gains approximately 0.5 to 1 perplexity point over the corresponding K-quant. At equivalent quality, it saves 20-30% of VRAM. These are rough orders of magnitude—your results will depend on the model and calibration dataset.

#Measured VRAM savings by model size

The figures below are rough estimates for recent dense models (Qwen 3.5, Granite 4.2), with a 4k context and no Flash Attention. Add 10-20% headroom for the KV cache and runtime overhead.

7B — Q4_K_M: 4.4 GB
TurboQuant 3.0 bpw: ~2.9 GB. Limited benefit; the 7B already fits everywhere.
13B — Q4_K_M: 7.8 GB
TurboQuant 3.0 bpw: ~5.1 GB. Enables 16k+ context on 8 GB of VRAM.
32B — Q4_K_M: 19 GB
TurboQuant 3.0 bpw: ~13 GB. Fits comfortably on RTX 4080 16 GB with an 8k context.
70B — Q4_K_M: 40 GB
TurboQuant 2.5 bpw: ~23 GB. Enables 70B on 1×RTX 4090 24 GB or 1×3090 24 GB.
120B MoE — Q4_K_M: ~70 GB
TurboQuant 2.7 bpw: ~42 GB. Feasible on 2×3090 or a 64 GB Mac Studio.
!
VRAM ≠ disk
A 23 GB TurboQuant file can consume 26–28 GB in practice once loaded: decompression, activation buffers, and KV cache. Always allow for 15–20% headroom on your total VRAM.

#Impact on quality

On the usual benchmarks (MMLU, HellaSwag, HumanEval), a 3.0 bpw TurboQuant version of a 70B remains within 0.5–1.5 points of FP16. On multi-step reasoning and long-form code, clearer differences begin to appear. The test that separates the methods most clearly: a long Python code excerpt with cross-dependencies.

General knowledge
Almost imperceptible. If you ask general-knowledge questions, you won't notice the difference.
Short reasoning
Minor loss. The chains of thought remain coherent over 5–10 steps.
Long reasoning (math/proofs)
Noticeable degradation. The model may skip a step or lose track after 20+ reasoning turns. Favor 3.5+ bpw if that is your use case.
Code
Sensitive to aggressive quantization. Below 3.0 bpw, expect more subtle bugs (wrong indexes, reversed conditions). Stick with Q4_K_M or TurboQuant ≥3.5 bpw for Aider/Continue.
Rare languages
French works well down to 2.5 bpw. Languages that are underrepresented in the calibration corpus suffer more.
→
The calibration dataset matters
A TurboQuant produced with a French dataset will perform better in French than one calibrated on standard English C4. If your use case is targeted, look at the community variants on Hugging Face—there’s often a more suitable “-fr” or “-code” version.

#Hardware and software requirements

Quantizing a model yourself is still a demanding task. Downloading a pre-quantized version from Hugging Face is almost always simpler. If you still want to produce your own version:

GPU for calibration
You must be able to load the model in FP16 or BF16 during analysis. For a 70B, that means 2×A100 80 GB or a 192 GB Mac Studio M2 Ultra. For a 32B, a RTX 4090 24 GB is sufficient for partial offload.
Base model
The unquantized weights (safetensors) downloaded from the original Hugging Face repo. Allow 140 GB for a 70B FP16 model.
Calibration dataset
256 to 1024 representative samples of your use case. French Wikipedia, code, and your own prompts. Avoid generic datasets if you have a specific domain.
Target runtime
Decide in advance: ExLlamaV3 (fastest on Nvidia), llama.cpp with the turbo backend, or a specific fork. The file format differs between runtimes.
Python 3.10+ and CUDA 12+
Standard toolchain. On AMD, ROCm 6.x works with llama.cpp but not yet with all turbo forks.

#Workflow for quantizing it yourself

The typical process, from an FP16 Hugging Face model to an executable file in Ollama or llama.cpp.

  1. 01
    Retrieve the original weights
    Clone the Hugging Face repo for the unquantized model. Allow at least an hour for a 70B model with a good connection. Use huggingface-cli download so you can resume an interrupted download.
  2. 02
    Preparing the calibration dataset
    Build a JSONL file with 256–1024 prompts. Diversity matters more than quantity: code, French prose, dialogues, technical questions. A 5–20 MB file is more than enough.
  3. 03
    Start calibration
    The tool (for example, the quantize script provided by the selected TurboQuant project) loads the model, runs each prompt forward, and collects layer-by-layer activation statistics. Allow 1-4h on a 70B, depending on the GPU.
  4. 04
    Quantify and export
    Bit allocation is calculated from the statistics, then each tensor is encoded. Output: a .safetensors or .gguf file, depending on the runtime. Allow an additional 30 minutes to 2 hours.
  5. 05
    Convert to the runtime format
    For Ollama: create a Modelfile pointing to the quantized file, then ollama create monmodele -f Modelfile. For llama.cpp: the main binary accepts the .gguf directly.
  6. 06
    Validate the quality
    Run a small test suite: 20–30 prompts covering your use cases. Compare side by side with the reference GGUF Q4_K_M version. If the degradation is too severe, repeat with more bits or a better dataset.
Ollama loading example
# Une fois le .gguf TurboQuant produit
cat > Modelfile <<EOF
FROM ./glm-4.7-turboquant-3.0bpw.gguf
PARAMETER num_ctx 8192
PARAMETER temperature 0.7
EOF

ollama create glm-4.7-turbo -f Modelfile
ollama run glm-4.7-turbo "Explique le théorème de Bayes en 3 phrases."
→
Download first, quantize afterward
Before starting 6 hours of calibration, search Hugging Face: for popular models (GLM 4.7, Qwen3, DeepSeek), there is almost always a community TurboQuant or EXL3 variant ready to use. Filter by “turbo,” “exl3,” or “3.0bpw.”

#Common pitfalls and troubleshooting

The runtime does not recognize the file
TurboQuant is not a single standard format. Make sure you're loading the file with the runtime that corresponds to the method used (ExLlamaV3 for EXL3, a recent version of llama.cpp for turbo GGUFs).
VRAM exceeded during inference even though the file fit
You’re forgetting the KV cache. With a 32k context, the KV cache can take up an additional 4–8 GB. Reduce num_ctx or enable Flash Attention with OLLAMA_FLASH_ATTENTION=1.
Catastrophic quality for your use case
The calibration dataset does not cover your domain. Re-run calibration with 100–200 representative prompts. This is by far the most powerful lever.
The quantized model is slower than Q4_K_M
Normal on some architectures: turbo decompression has overhead. On older cards (RTX 20xx), the VRAM gain may come at the cost of tokens/sec. Measure before adopting it.
Differences between runs
Calibration introduces nondeterminism. Two quantizations of the same model with the same dataset may differ slightly in quality. Run several times and keep the best one.
!
Check the license
Quantizing a model does not change its original license. A Gemma 4 remains under Apache 2.0, a DeepSeek remains under MIT, and a Codestral 22B remains prohibited in production even when quantized. If you redistribute your TurboQuant version on Hugging Face, keep the LICENSE file and attribution.

#Go further

TurboQuant is one of several tools for pushing the limit of what fits locally. A few complementary options:

Choose your quantization (Q4, Q5, Q8, FP16)
To properly position TurboQuant against classic GGUF K-quants and choose on a case-by-case basis.
Local LLM fine-tuning: LoRA and QLoRA
If you want to go beyond quantization and adapt a model to your domain—often a better lever than more aggressive quantization.
Choose your GPU for local AI
To decide which card to target, knowing that VRAM remains the number-one constraint, even with TurboQuant.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.