Beginner 11 minMacBook Air

Which LLM on MacBook Air M4 (16 / 24 / 32 GB) ?

Direct response

The MacBook Air M4 comfortably runs 8- to 12-billion-parameter models on 16 GB, a 24B in Q4 on 24 GB, and a 30B mixture-of-experts model on 32 GB. Its 120 GB/s bandwidth caps an 8B in Q4 at approximately 24 tokens per second; non-expandable memory determines the model size.

The M4 is the first Air to start at 16 GB and go up to 32 GB. This page explains which memory size to choose, calculates the maximum speed the chip can reach, compares it with the Air M5, the MacBook Pro, and the Mac mini, and shows where the Air reaches its limits.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#MacBook Air M4: 120 GB/s and up to 32 GB of memory

The MacBook Air M4 is the first Air to start with 16 GB of unified memory and support up to 32 GB, with 120 GB/s of bandwidth. For an LLM, these two figures matter more than the core count: memory determines the size of the model that can be loaded, while bandwidth determines its maximum speed. At 120 GB/s, an 8-billion-parameter model in Q4 cannot theoretically exceed 24 tokens per second; real-world speed is lower.

Chip
Apple M4: 10-core CPU (4 performance, 6 efficiency), 8- or 10-core GPU, 16-core Neural Engine, according to the Apple technical specifications.
Memory
16 GB standard, configurable to 24 or 32 GB. It's soldered: you choose the configuration when purchasing.
Bandwidth
120 GB/s, according to the Apple specification sheet, or 20% more than the Air M3's 100 GB/s.
Neural Engine
Apple reports up to 38,000 billion operations per second for the M4. Common engines (Ollama, llama.cpp) run mainly on the GPU through Metal, so this figure does not translate into tokens per second.
Cooling
Fanless: heat is dissipated through the chassis, which limits very long workloads.

There is no separate VRAM on a Mac: the GPU uses unified memory. When you search for "MacBook Air M4 VRAM," the practical answer is total memory minus the portion reserved for the system, within the limit macOS allows the GPU to use. This limit can be adjusted; the site's macOS optimization guide explains how and what the risks are.

#16, 24, or 32 GB: choosing your memory

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The model, its context, and the system share memory. The margins below are conservative estimates to cross-check against memory pressure in Activity Monitor: roughly 11 to 12 GB for a model on 16 GB, 17 to 19 GB on 24 GB, and 24 to 26 GB on 32 GB. Model sizes follow the site's reference points: 8B is about 5 GB, 14B about 9 GB, 24B about 14 GB, and 32B about 19 to 20 GB in Q4.

MacBook Air M4: model class by memory (Q4_K_M)
MemoryEstimated marginComfortable tierOn the limit
16 GB≈ 11 to 12 GB8–9B with medium context, 12B in Q414B in Q4 with a short context
24 GB≈ 17 to 19 GB12–14B comfortable, 24B in Q4 (≈ 14 GB)27B in Q4 (≈ 16 GB) with reduced context
32 GB≈ 24 to 26 GB24B in Q5–Q6, 27B in Q4, 30B MoE modelsDense 32B in Q4 (≈ 19-20 GB) with a long context
i
The useful tier depends on model generation
Mixture-of-experts (MoE) models change the computation: a 30B model with 3B active parameters takes about 19 GB of memory but reads only a small fraction of the weights for each token. On 32 GB, it runs much faster than a dense model with the same parameter count. On 16 GB, it will not load: memory remains the first filter.

The question, then, is what size you want to target. If you limit yourself to models with 8 to 12 billion parameters, 16 GB is enough. For a 24B or for a RAG with an embedder and reranker running in parallel, 24 GB is the first comfortable tier. At 32 GB, you mainly gain long context and 30B MoE models.

#Maximum speed at 120 GB/s

The principle: each generated token forces the active weights to be read. The ceiling is bandwidth divided by weight in gigabytes. The table applies this formula to the M4 (120 GB/s). These are calculated ceilings, not measurements: actual speed depends on the engine, quantization, context, and chip temperature.

Theoretical generation ceiling on Air M4 (120 GB/s, active weights in Q4): calculation, not measurement
ModelActive weightsComputeCap
3B≈ 2 GB120 ÷ 2≈ 60 tok/s
8B≈ 5 GB120 ÷ 5≈ 24 tok/s
9B≈ 6 GB120 ÷ 6≈ 20 tok/s
14B≈ 9 GB120 ÷ 9≈ 13 tok/s
24B≈ 14 GB120 ÷ 14≈ 8 to 9 tok/s
32B dense≈ 19 to 20 GB120 ÷ 19,5≈ 6 tok/s

Two practical consequences. An 8B model in Q4 will not exceed about 24 tokens per second on this chip: a claimed figure of 34 tokens per second for this model is impossible without another technique. And a dense 24B model remains around 8 tokens per second at best, roughly the speed of attentive reading: comfortable for a composed text, slow for an agent that chains calls. For real-world measurements, see /benchmarks and third parties that specify the chip and quantization.

#M4 or M5: does the next generation change the picture?

Apple now offers an M5 MacBook Air. Its specifications list 153 GB/s of bandwidth, versus 120 GB/s for the M4, or about 28% more, with the same 16, 24, or 32 GB memory options. For an LLM limited by bandwidth, the theoretical maximum speed gap is of the same order: an 8B in Q4 would go from about 24 to about 30 tokens per second at the ceiling.

Air M4 and Air M5: the numbers that matter for an LLM, according to Apple
CriterionAir M4Air M5
Bandwidth120 GB/s153 GB/s
Memory16, 24, or 32 GB16, 24, or 32 GB
Theoretical 8B Q4 ceiling≈ 24 tok/s≈ 30 tok/s

Memory itself doesn't change: a 16 GB M5 Air doesn't open more models than a 16 GB M4 Air. If you find an M4 configured with 24 or 32 GB, it's better than a 16 GB M5 for model size; with equal memory, the M5 wins on speed. Compare current in-store prices; they change too much to list here.

#Install and choose your engine

  1. 01
    Install Ollama
    Download it from ollama.com or with Homebrew. The macOS installation guide covers automatic startup.
  2. 02
    Run a model based on available memory
    9B on 16 GB, 24B in Q4 on 24 GB, a 30B MoE on 32 GB. The first launch downloads the model.
  3. 03
    Check the GPU
    ollama ps affiche la répartition CPU/GPU dans la colonne PROCESSOR : tout sur le GPU est l'état attendu.
  4. 04
    Adjust the context
    Ollama sets the default context to 4,000 tokens with less than 24 GiB of memory. Increase it as needed; the KV cache grows with it.
Terminal
brew install --cask ollama
ollama run qwen3.5:9b
ollama ps

MLX, Apple's framework, uses the Mac's unified memory without copying data between the CPU and GPU; LM Studio offers it as MLX models. Both approaches work: Ollama for simplicity and a local API, MLX and LM Studio for exploring Apple formats. The MLX versus llama.cpp guide compares the two approaches, so there is no need to repeat its conclusions here.

#Long context and GPU memory limit

The model is only part of the memory footprint: each context token adds an entry to the KV cache, and the memory actually available to the GPU is capped by macOS. Ollama uses flash attention automatically when the engine and hardware support it, and lets you quantize the KV cache with the OLLAMA_KV_CACHE_TYPE variable (f16 by default). On 16 GB, this setting often separates a short context from a context of several thousand tokens.

On the MLX side, the mlx-lm repository indicates that a model that fits in memory can often be accelerated by raising the system's wired memory limit via a sysctl setting. This is a system change: it requires administrator privileges, may not survive a reboot, and can destabilize the machine if the value is too high. The macOS optimization guide details the procedure and the reasonable value for the available memory; do not change it without reading the guide.

Terminal
export OLLAMA_KV_CACHE_TYPE=q8_0
export OLLAMA_FLASH_ATTENTION=1
ollama serve

#M4 Air, MacBook Pro, or Mac mini: which one for which use case

The M4 Air wins on lightness and silence, but it has no fan: sustained workloads reduce the chip’s clock speed. The MacBook Pro M4 Pro or Max adds fans, 273 to 546 GB/s of bandwidth, and up to 128 GB of memory according to the Apple spec sheet; the Mac mini M4 provides a stationary, actively cooled machine.

Choose based on your use case
Your needsSuitable machineWhy
Chat, summarization, light coding, mobilityAir M4 16 or 24 GBQuiet operation, battery life, 8B to 14B models
24B to 32B models, long contextAir M4 32 GB or Mac miniSufficient memory; Mac mini if the machine stays at the desk
Long-running jobs, multiple modelsMacBook Pro M4 Pro or Max, Mac miniActive cooling, higher bandwidth
70B modelsMacBook Pro Max or Mac StudioA 70B weighs approximately 40 GB in Q4: out of reach for an Air

#Limitations to know

Long-running workload
Without a fan, continuous generation for several dozen minutes eventually reduces the clock speed. For chat, the effect goes unnoticed; for indexing or batch processing, a fan-cooled machine is preferable.
No 70B
A 70B model weighs approximately 40 GB in Q4: more than the Air's maximum 32 GB.
The Neural Engine does not accelerate Ollama
Common engines use the Metal GPU. The NPU only comes into play for certain workloads with Core ML.
Storage
Each model takes up several gigabytes in Ollama's model directory: a 256 GB SSD fills up quickly with several 5 to 15 GB models.

#Frequently asked questions

FAQ
Is 16 GB on the MacBook Air M4 Enough for a Local LLM?+
Yes, for 8- to 12-billion-parameter models in Q4, which weigh 5 to 8 GB and leave room for the system and context. A 14-billion-parameter model works with a short context. For 24B or heavy RAG workloads, target 24 or 32 GB, since the memory cannot be expanded.
How much VRAM does an M4 MacBook Air have?+
No separate VRAM: the GPU shares unified memory with the processor, totaling 16, 24, or 32 GB depending on the configuration. macOS caps the amount the GPU can reserve, so the loaded model cannot use all the memory. The macOS optimization guide explains how to raise this limit.
What speed can you expect from an 8B model on a MacBook Air M4?+
The theoretical ceiling is 120 GB/s divided by about 5 GB, or at most 24 tokens per second for an 8B in Q4. Actual speed is lower and depends on the engine and context. A reported throughput significantly above this ceiling should raise suspicion.
Can you run a 30-billion-parameter model on an M4 Air?+
Only on 32 GB, and preferably with a mixture-of-experts model with few active parameters. A 30B weighs about 19 GB in Q4: it fits within the 24–26 GB margin, but a dense 30B remains slow, at around 6 tokens per second at best. It does not load on 16 GB.
Should you wait for the Air M5 for local AI?+
The M5 offers 153 GB/s versus 120 GB/s, or about 28% more bandwidth, but the same 16, 24, or 32 GB options. If you find an M4 with more memory than an M5 at the same budget, choose the memory: it determines the model size.
Does the MacBook Air M4 heat up with an LLM?+
It has no fan, so it heats up and throttles under sustained load. The effect is minor during short exchanges. For long-running tasks, connect the power adapter, elevate the machine, and choose a smaller model. For continuous workloads, prefer a Mac mini or MacBook Pro.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.