MLX: LLMs optimized for Apple Silicon (M1-M4)

Run a MLX Apple Silicon LLM on a recent Mac, it means leveraging the unified memory of M1 to M4 chips to run 7- to 235-billion-parameter open-weights models locally. The MLX framework, released by Apple Research in late 2023, directly leverages the Neural Engine and integrated GPU through Metal, without going through CUDA. This guide explains how MLX works, compatible models, realistic VRAM requirements by quantization, observed speeds on the M3 Max and M4 Pro, reference benchmarks, and the strongest use cases for self-hosted deployment on macOS.

Why MLX changes the game on Mac

MLX is an array-computing library inspired by NumPy and PyTorch, designed specifically for the unified memory of Apple Silicon. Unlike llama.cpp which uses Metal as an optional backend, MLX treats CPU and GPU memory as a single pool: no copying is required between them. The official repository ml-explore/mlx on GitHub publishes the sources under the MIT license, and the associated project mlx-examples provides ready-to-use implementations for Llama, Mistral, Phi, and Qwen.

Key takeaways:

To compare Mac and Linux approaches, see llama.cpp and GGUF as well as our comparison MLX vs llama.cpp.

Required VRAM and supported quantizations

MLX exposes several quantization levels. Here are the approximate figures, starting from the FP16 numbers and applying the usual ratios (Q8 ≈ 50%, Q4 ≈ 25%, Q3 ≈ 20%).

Models that run comfortably on a 128 GB MacBook Pro M3 Max or 128 GB M4 Max:

Larger models reserved for the 192 GB or 512 GB Mac Studio M3 Ultra:

At very large scale, DeepSeek V3 671B (Q4 VRAM ~400 GB, 128k context) remains out of reach even for a 512 GB Mac Studio in Q4, but becomes feasible in Q2/Q3 with a little disk offloading. Full details on DeepSeek V3 671B.

Measured performance: tokens per second

The figures below come from the official documentation mlx-lm and reproducible community observations on Hugging Face. All speeds are estimates and depend on the prompt, context, and quantization level.

MoE models such as Qwen3-Coder-Next 80B-A3B (Q4 VRAM ~48 GB) particularly benefit from unified memory: only 3B active parameters pass through the compute units at each token, pushing speed above 25 tokens/sec on M4 Max according to user reports (to be confirmed). See Qwen3-Coder-Next 80B-A3B.

Licensing and compatibility with MLX

MLX imposes no restrictions beyond the original model's license. A few typical cases:

For production use, favor Apache 2.0 or MIT. See our guide open-weights LLM licenses for the details of the clauses.

Reference benchmarks for MLX-compatible models

The scores reported here come from official technical specifications published on Hugging Face and model cards from the vendors.

For a detailed look at the best-performing coding models, see best LLM for coding and for reasoning best reasoning LLM.

Installation and practical workflow

Installation takes three commands after preparing a Python 3.10+ environment:

pip install mlx mlx-lm
python -m mlx_lm.convert --hf-path meta-llama/Llama-3.3-70B-Instruct -q
python -m mlx_lm.generate --model ./mlx_model --prompt "Bonjour"

The conversion -q applies 4-bit quantization by default. The converted checkpoints are published directly on Hugging Face mlx-community, which avoids the local conversion step for popular models.

Use cases where MLX particularly shines:

To compare the MacBook Pro and Mac Studio for your target model, see LLM for MacBook Pro.

FAQ

Q: Which Mac should you choose to run Llama 3.3 70B locally?

A 64 GB MacBook Pro M3 Max or M4 Max can run Llama 3.3 70B Instruct in 4-bit quantization (VRAM ~40 GB). With 48 GB, it's tight but possible in Q3. For long contexts or batches, target 128 GB. The Mac Studio M3 Ultra 192 GB remains the most comfortable target for this model.

Q: Is MLX faster than llama.cpp on Mac?

MLX generally gains 10 to 30% on long-form generation thanks to native unified memory and the graph compiler. llama.cpp remains competitive on prefill and offers a broader ecosystem (built-in OpenAI-compatible server, more varied GGUF quantizations). Both can coexist: MLX for development, llama.cpp for headless deployment.

Q: Can you do LoRA fine-tuning with MLX?

Yes. The subproject mlx-examples/lora lets you train LoRA and QLoRA adapters on Llama, Mistral, and Qwen. On an M3 Max with 64 GB, LoRA fine-tuning of Mistral Small 4 4-bit remains out of reach, but 7B to 13B models can be trained in a few hours.

Q: Does MLX support MoE models such as Mixtral or Qwen 3 235B?

Yes. Mixtral 8x22B Instruct et Qwen 3 235B-A22B are supported via mlx-lm. The MoE architecture is particularly well suited to unified memory: inactive experts reside in RAM without penalizing latency. On a 192 GB Mac Studio M3 Ultra, Qwen 3 235B-A22B in 4-bit remains usable with a reasonable context.

Q: Which models are not viable on Apple Silicon?

Models >500B at full precision remain out of reach even for the 512 GB Mac Studio: DeepSeek V4 Pro 1.6T (Q4 VRAM ~960 GB), MiMo V2.5 Pro (Q4 VRAM ~595 GB), Kimi K2.6 (Q4 VRAM ~600 GB). With aggressive Q2 or Q3 quantization and disk offload, some become experimentally testable, but speed drops drastically.

Q: Does MLX support vision-language models?

Yes, partially. Qwen 2.5 VL 72B et LLaVA-OneVision 72B have community MLX ports on Hugging Face. Coverage is more limited than for pure text LLMs, but the project mlx-vlm community-maintained, closes the gap on common multimodal architectures.

Conclusion

Un MLX Apple Silicon LLM turns a recent Mac into a serious inference workstation: Llama 3.3 70B, Qwen 2.5 72B, or gpt-oss 120B run in 4-bit on a 64-128 GB MacBook Pro, and MoE models such as Qwen 3 235B-A22B become accessible on a 192 GB Mac Studio. To choose the right model for your exact hardware configuration, use our configurator or browse the full catalog of the 249 indexed models.

Article published on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.