MLX: LLMs optimized for Apple Silicon (M1-M4)
Run a MLX Apple Silicon LLM on a recent Mac, it means leveraging the unified memory of M1 to M4 chips to run 7- to 235-billion-parameter open-weights models locally. The MLX framework, released by Apple Research in late 2023, directly leverages the Neural Engine and integrated GPU through Metal, without going through CUDA. This guide explains how MLX works, compatible models, realistic VRAM requirements by quantization, observed speeds on the M3 Max and M4 Pro, reference benchmarks, and the strongest use cases for self-hosted deployment on macOS.
Why MLX changes the game on Mac
MLX is an array-computing library inspired by NumPy and PyTorch, designed specifically for the unified memory of Apple Silicon. Unlike llama.cpp which uses Metal as an optional backend, MLX treats CPU and GPU memory as a single pool: no copying is required between them. The official repository ml-explore/mlx on GitHub publishes the sources under the MIT license, and the associated project mlx-examples provides ready-to-use implementations for Llama, Mistral, Phi, and Qwen.
Key takeaways:
- Unified memory : a Mac Studio M3 Ultra with 192 GB exposes all 192 GB to the GPU. A PC with a RTX 4090 is capped at 24 GB of GDDR6X VRAM.
- Native quantization : MLX supports int4, int8, and bfloat16 via
mlx_lm.convert. - Lazy evaluation : the computation graphs are materialized only when called
mx.eval(), which reduces consumption. - Hugging Face compatibility : the checkpoints
safetensorsstandards are convertible with a command.
To compare Mac and Linux approaches, see llama.cpp and GGUF as well as our comparison MLX vs llama.cpp.
Required VRAM and supported quantizations
MLX exposes several quantization levels. Here are the approximate figures, starting from the FP16 numbers and applying the usual ratios (Q8 ≈ 50%, Q4 ≈ 25%, Q3 ≈ 20%).
Models that run comfortably on a 128 GB MacBook Pro M3 Max or 128 GB M4 Max:
- gpt-oss 120B : Q4 VRAM ~70 GB, 128k-token context, Apache 2.0 license. See gpt-oss 120B.
- Llama 3.3 70B Instruct : ~40 GB Q4 VRAM, 128k context. Reference for agentic workflows on Mac, details on Llama 3.3 70B.
- Qwen 2.5 72B Instruct : ~42 GB Q4 VRAM, 131k context. Complete profile on Qwen 2.5 72B.
- DeepSeek R1 Distill Llama 70B : Q4 VRAM ~40 GB. The best-known reasoning distillate, details on DeepSeek R1 Distill 70B.
Larger models reserved for the 192 GB or 512 GB Mac Studio M3 Ultra:
- Mixtral 8x22B Instruct : ~82 GB Q4 VRAM, 64k context, Apache 2.0. See Mixtral 8x22B.
- Mistral Small 4 (119B): Q4 VRAM ~72 GB, 256k context. Specs: Mistral Small 4.
- Qwen 3 235B-A22B : Q4 VRAM ~142 GB, 131k context. MoE architecture that activates only 22B parameters, listing Qwen 3 235B-A22B.
- Qwen 3.5 122B-A10B : Q4 VRAM ~73 GB, 262k context. See Qwen 3.5 122B-A10B.
At very large scale, DeepSeek V3 671B (Q4 VRAM ~400 GB, 128k context) remains out of reach even for a 512 GB Mac Studio in Q4, but becomes feasible in Q2/Q3 with a little disk offloading. Full details on DeepSeek V3 671B.
Measured performance: tokens per second
The figures below come from the official documentation mlx-lm and reproducible community observations on Hugging Face. All speeds are estimates and depend on the prompt, context, and quantization level.
- Llama 3.1 70B in 4-bit on an M3 Max 128 GB: ~10 to 12 tokens/sec during generation, ~150 tokens/sec prefill (estimated). Specs: Llama 3.1 70B.
- Qwen 2.5 72B in 4-bit on M4 Max: ~9 to 11 tokens/sec (to be confirmed depending on the SoC revision).
- gpt-oss 120B in 4-bit on M3 Ultra 192 GB: ~14 tokens/sec (estimated from Hugging Face benchmarks).
- Mixtral 8x22B in 4-bit on M3 Ultra: ~18 tokens/sec thanks to MoE activation.
- DeepSeek R1 Distill Llama 70B in 4-bit on an M3 Max: ~10 tokens/sec, including reasoning traces.
MoE models such as Qwen3-Coder-Next 80B-A3B (Q4 VRAM ~48 GB) particularly benefit from unified memory: only 3B active parameters pass through the compute units at each token, pushing speed above 25 tokens/sec on M4 Max according to user reports (to be confirmed). See Qwen3-Coder-Next 80B-A3B.
Licensing and compatibility with MLX
MLX imposes no restrictions beyond the original model's license. A few typical cases:
- Apache 2.0 (Mixtral 8x22B, Mistral Small 4, gpt-oss 120B, Qwen 3 235B-A22B, Qwen 3.5 397B-A17B): commercial use without restrictions.
- MIT (DeepSeek R1 671B, GLM-5.1, Ling 2.6 1T): highly permissive; attribution required.
- Llama Community (Llama 3.1 70B, Llama 3.1 405B, Llama 4 Scout 109B): permits commercial use below the threshold of 700 million monthly active users.
- Qwen License (Qwen 2.5 72B, Qwen 2.5 VL 72B): specific conditions for large platforms.
- CC-BY-NC 4.0 (Command R+ 104B): noncommercial use only.
For production use, favor Apache 2.0 or MIT. See our guide open-weights LLM licenses for the details of the clauses.
Reference benchmarks for MLX-compatible models
The scores reported here come from official technical specifications published on Hugging Face and model cards from the vendors.
- MMLU : Llama 3.3 70B reaches ~86, Qwen 2.5 72B ~85, gpt-oss 120B ~88 (estimated).
- HumanEval : Qwen3-Coder-Next 80B-A3B targets >85% pass@1, DeepSeek R1 Distill Llama 70B ~78% (to be confirmed).
- AIME 2024 : DeepSeek R1 671B reaches 79.8% according to the arXiv paper DeepSeek R1. The 70B distill achieves ~70%.
- GPQA Diamond : DeepSeek R1 ~71%, Llama 3.3 70B ~50% (estimated).
For a detailed look at the best-performing coding models, see best LLM for coding and for reasoning best reasoning LLM.
Installation and practical workflow
Installation takes three commands after preparing a Python 3.10+ environment:
pip install mlx mlx-lm
python -m mlx_lm.convert --hf-path meta-llama/Llama-3.3-70B-Instruct -q
python -m mlx_lm.generate --model ./mlx_model --prompt "Bonjour"
The conversion -q applies 4-bit quantization by default. The converted checkpoints are published directly on Hugging Face mlx-community, which avoids the local conversion step for popular models.
Use cases where MLX particularly shines:
- Independent Mac developers : a 64 GB MacBook Pro M3 Max is enough to run Llama 3.3 70B in Q3 or Qwen 2.5 72B in Q3.
- Local RAG workflows : unified memory lets you keep a long context (128k or even 256k tokens) without swapping.
- Private coding agents : Qwen3-Coder-Next 80B-A3B in 4-bit MoE remains smooth for IDE completion.
- Academic research : MLX integrates
mlx.optimizersto perform LoRA fine-tuning directly on Apple Silicon (see LoRA examples mlx-examples).
To compare the MacBook Pro and Mac Studio for your target model, see LLM for MacBook Pro.
FAQ
Q: Which Mac should you choose to run Llama 3.3 70B locally?
A 64 GB MacBook Pro M3 Max or M4 Max can run Llama 3.3 70B Instruct in 4-bit quantization (VRAM ~40 GB). With 48 GB, it's tight but possible in Q3. For long contexts or batches, target 128 GB. The Mac Studio M3 Ultra 192 GB remains the most comfortable target for this model.
Q: Is MLX faster than llama.cpp on Mac?
MLX generally gains 10 to 30% on long-form generation thanks to native unified memory and the graph compiler. llama.cpp remains competitive on prefill and offers a broader ecosystem (built-in OpenAI-compatible server, more varied GGUF quantizations). Both can coexist: MLX for development, llama.cpp for headless deployment.
Q: Can you do LoRA fine-tuning with MLX?
Yes. The subproject mlx-examples/lora lets you train LoRA and QLoRA adapters on Llama, Mistral, and Qwen. On an M3 Max with 64 GB, LoRA fine-tuning of Mistral Small 4 4-bit remains out of reach, but 7B to 13B models can be trained in a few hours.
Q: Does MLX support MoE models such as Mixtral or Qwen 3 235B?
Yes. Mixtral 8x22B Instruct et Qwen 3 235B-A22B are supported via mlx-lm. The MoE architecture is particularly well suited to unified memory: inactive experts reside in RAM without penalizing latency. On a 192 GB Mac Studio M3 Ultra, Qwen 3 235B-A22B in 4-bit remains usable with a reasonable context.
Q: Which models are not viable on Apple Silicon?
Models >500B at full precision remain out of reach even for the 512 GB Mac Studio: DeepSeek V4 Pro 1.6T (Q4 VRAM ~960 GB), MiMo V2.5 Pro (Q4 VRAM ~595 GB), Kimi K2.6 (Q4 VRAM ~600 GB). With aggressive Q2 or Q3 quantization and disk offload, some become experimentally testable, but speed drops drastically.
Q: Does MLX support vision-language models?
Yes, partially. Qwen 2.5 VL 72B et LLaVA-OneVision 72B have community MLX ports on Hugging Face. Coverage is more limited than for pure text LLMs, but the project mlx-vlm community-maintained, closes the gap on common multimodal architectures.
Conclusion
Un MLX Apple Silicon LLM turns a recent Mac into a serious inference workstation: Llama 3.3 70B, Qwen 2.5 72B, or gpt-oss 120B run in 4-bit on a 64-128 GB MacBook Pro, and MoE models such as Qwen 3 235B-A22B become accessible on a 192 GB Mac Studio. To choose the right model for your exact hardware configuration, use our configurator or browse the full catalog of the 249 indexed models.