Which LLM on MacBook Air M4 (16 / 24 / 32 GB) ?
The MacBook Air M4 comfortably runs 8- to 12-billion-parameter models on 16 GB, a 24B in Q4 on 24 GB, and a 30B mixture-of-experts model on 32 GB. Its 120 GB/s bandwidth caps an 8B in Q4 at approximately 24 tokens per second; non-expandable memory determines the model size.
The M4 is the first Air to start at 16 GB and go up to 32 GB. This page explains which memory size to choose, calculates the maximum speed the chip can reach, compares it with the Air M5, the MacBook Pro, and the Mac mini, and shows where the Air reaches its limits.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#MacBook Air M4: 120 GB/s and up to 32 GB of memory
The MacBook Air M4 is the first Air to start with 16 GB of unified memory and support up to 32 GB, with 120 GB/s of bandwidth. For an LLM, these two figures matter more than the core count: memory determines the size of the model that can be loaded, while bandwidth determines its maximum speed. At 120 GB/s, an 8-billion-parameter model in Q4 cannot theoretically exceed 24 tokens per second; real-world speed is lower.
- Chip
- Apple M4: 10-core CPU (4 performance, 6 efficiency), 8- or 10-core GPU, 16-core Neural Engine, according to the Apple technical specifications.
- Memory
- 16 GB standard, configurable to 24 or 32 GB. It's soldered: you choose the configuration when purchasing.
- Bandwidth
- 120 GB/s, according to the Apple specification sheet, or 20% more than the Air M3's 100 GB/s.
- Neural Engine
- Apple reports up to 38,000 billion operations per second for the M4. Common engines (Ollama, llama.cpp) run mainly on the GPU through Metal, so this figure does not translate into tokens per second.
- Cooling
- Fanless: heat is dissipated through the chassis, which limits very long workloads.
There is no separate VRAM on a Mac: the GPU uses unified memory. When you search for "MacBook Air M4 VRAM," the practical answer is total memory minus the portion reserved for the system, within the limit macOS allows the GPU to use. This limit can be adjusted; the site's macOS optimization guide explains how and what the risks are.
#16, 24, or 32 GB: choosing your memory
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
The model, its context, and the system share memory. The margins below are conservative estimates to cross-check against memory pressure in Activity Monitor: roughly 11 to 12 GB for a model on 16 GB, 17 to 19 GB on 24 GB, and 24 to 26 GB on 32 GB. Model sizes follow the site's reference points: 8B is about 5 GB, 14B about 9 GB, 24B about 14 GB, and 32B about 19 to 20 GB in Q4.
| Memory | Estimated margin | Comfortable tier | On the limit |
|---|---|---|---|
| 16 GB | ≈ 11 to 12 GB | 8–9B with medium context, 12B in Q4 | 14B in Q4 with a short context |
| 24 GB | ≈ 17 to 19 GB | 12–14B comfortable, 24B in Q4 (≈ 14 GB) | 27B in Q4 (≈ 16 GB) with reduced context |
| 32 GB | ≈ 24 to 26 GB | 24B in Q5–Q6, 27B in Q4, 30B MoE models | Dense 32B in Q4 (≈ 19-20 GB) with a long context |
The question, then, is what size you want to target. If you limit yourself to models with 8 to 12 billion parameters, 16 GB is enough. For a 24B or for a RAG with an embedder and reranker running in parallel, 24 GB is the first comfortable tier. At 32 GB, you mainly gain long context and 30B MoE models.
#Maximum speed at 120 GB/s
The principle: each generated token forces the active weights to be read. The ceiling is bandwidth divided by weight in gigabytes. The table applies this formula to the M4 (120 GB/s). These are calculated ceilings, not measurements: actual speed depends on the engine, quantization, context, and chip temperature.
| Model | Active weights | Compute | Cap |
|---|---|---|---|
| 3B | ≈ 2 GB | 120 ÷ 2 | ≈ 60 tok/s |
| 8B | ≈ 5 GB | 120 ÷ 5 | ≈ 24 tok/s |
| 9B | ≈ 6 GB | 120 ÷ 6 | ≈ 20 tok/s |
| 14B | ≈ 9 GB | 120 ÷ 9 | ≈ 13 tok/s |
| 24B | ≈ 14 GB | 120 ÷ 14 | ≈ 8 to 9 tok/s |
| 32B dense | ≈ 19 to 20 GB | 120 ÷ 19,5 | ≈ 6 tok/s |
Two practical consequences. An 8B model in Q4 will not exceed about 24 tokens per second on this chip: a claimed figure of 34 tokens per second for this model is impossible without another technique. And a dense 24B model remains around 8 tokens per second at best, roughly the speed of attentive reading: comfortable for a composed text, slow for an agent that chains calls. For real-world measurements, see /benchmarks and third parties that specify the chip and quantization.
#M4 or M5: does the next generation change the picture?
Apple now offers an M5 MacBook Air. Its specifications list 153 GB/s of bandwidth, versus 120 GB/s for the M4, or about 28% more, with the same 16, 24, or 32 GB memory options. For an LLM limited by bandwidth, the theoretical maximum speed gap is of the same order: an 8B in Q4 would go from about 24 to about 30 tokens per second at the ceiling.
| Criterion | Air M4 | Air M5 |
|---|---|---|
| Bandwidth | 120 GB/s | 153 GB/s |
| Memory | 16, 24, or 32 GB | 16, 24, or 32 GB |
| Theoretical 8B Q4 ceiling | ≈ 24 tok/s | ≈ 30 tok/s |
Memory itself doesn't change: a 16 GB M5 Air doesn't open more models than a 16 GB M4 Air. If you find an M4 configured with 24 or 32 GB, it's better than a 16 GB M5 for model size; with equal memory, the M5 wins on speed. Compare current in-store prices; they change too much to list here.
#Install and choose your engine
- 01Install OllamaDownload it from ollama.com or with Homebrew. The macOS installation guide covers automatic startup.
- 02Run a model based on available memory9B on 16 GB, 24B in Q4 on 24 GB, a 30B MoE on 32 GB. The first launch downloads the model.
- 03Check the GPUollama ps affiche la répartition CPU/GPU dans la colonne PROCESSOR : tout sur le GPU est l'état attendu.
- 04Adjust the contextOllama sets the default context to 4,000 tokens with less than 24 GiB of memory. Increase it as needed; the KV cache grows with it.
MLX, Apple's framework, uses the Mac's unified memory without copying data between the CPU and GPU; LM Studio offers it as MLX models. Both approaches work: Ollama for simplicity and a local API, MLX and LM Studio for exploring Apple formats. The MLX versus llama.cpp guide compares the two approaches, so there is no need to repeat its conclusions here.
#Long context and GPU memory limit
The model is only part of the memory footprint: each context token adds an entry to the KV cache, and the memory actually available to the GPU is capped by macOS. Ollama uses flash attention automatically when the engine and hardware support it, and lets you quantize the KV cache with the OLLAMA_KV_CACHE_TYPE variable (f16 by default). On 16 GB, this setting often separates a short context from a context of several thousand tokens.
On the MLX side, the mlx-lm repository indicates that a model that fits in memory can often be accelerated by raising the system's wired memory limit via a sysctl setting. This is a system change: it requires administrator privileges, may not survive a reboot, and can destabilize the machine if the value is too high. The macOS optimization guide details the procedure and the reasonable value for the available memory; do not change it without reading the guide.
#M4 Air, MacBook Pro, or Mac mini: which one for which use case
The M4 Air wins on lightness and silence, but it has no fan: sustained workloads reduce the chip’s clock speed. The MacBook Pro M4 Pro or Max adds fans, 273 to 546 GB/s of bandwidth, and up to 128 GB of memory according to the Apple spec sheet; the Mac mini M4 provides a stationary, actively cooled machine.
| Your needs | Suitable machine | Why |
|---|---|---|
| Chat, summarization, light coding, mobility | Air M4 16 or 24 GB | Quiet operation, battery life, 8B to 14B models |
| 24B to 32B models, long context | Air M4 32 GB or Mac mini | Sufficient memory; Mac mini if the machine stays at the desk |
| Long-running jobs, multiple models | MacBook Pro M4 Pro or Max, Mac mini | Active cooling, higher bandwidth |
| 70B models | MacBook Pro Max or Mac Studio | A 70B weighs approximately 40 GB in Q4: out of reach for an Air |
- MacBook Pro M4 Pro and Max: 273 to 546 GB/s
- Mac mini M4 and M4 Pro for an LLM
- MacBook Air M2: 100 GB/s, 24 GB maximum
- MLX vs. llama.cpp on Mac
- Optimize a Mac Apple Silicon for LLMs
- Which LLM for 16 GB of memory
- Source: Apple technical specifications for the MacBook Air M4
- Source: Apple technical specifications for the current MacBook Air
- Source: Apple technical specifications for the MacBook Pro M4 Pro and Max
- Source: Ollama documentation, context
#Limitations to know
- Long-running workload
- Without a fan, continuous generation for several dozen minutes eventually reduces the clock speed. For chat, the effect goes unnoticed; for indexing or batch processing, a fan-cooled machine is preferable.
- No 70B
- A 70B model weighs approximately 40 GB in Q4: more than the Air's maximum 32 GB.
- The Neural Engine does not accelerate Ollama
- Common engines use the Metal GPU. The NPU only comes into play for certain workloads with Core ML.
- Storage
- Each model takes up several gigabytes in Ollama's model directory: a 256 GB SSD fills up quickly with several 5 to 15 GB models.
#Frequently asked questions
Is 16 GB on the MacBook Air M4 Enough for a Local LLM?+
How much VRAM does an M4 MacBook Air have?+
What speed can you expect from an 8B model on a MacBook Air M4?+
Can you run a 30-billion-parameter model on an M4 Air?+
Should you wait for the Air M5 for local AI?+
Does the MacBook Air M4 heat up with an LLM?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.