Which LLM on a MacBook Air M2 (8 / 16 / 24 GB) ?
The MacBook Air M2 runs 3- to 4-billion-parameter models on 8 GB, 8- to 9-billion-parameter models on 16 GB, and up to a 24B model in Q4 on 24 GB. Its 100 GB/s bandwidth, 50% more than the M1 according to Apple, caps an 8B model in Q4 at about 20 tokens per second; memory determines what can load.
The M2 is the first Air chip to offer 24 GB and 100 GB/s. This page explains which class of model to choose based on 8, 16, or 24 GB, calculates the maximum physically possible speed, and indicates when to move to an M4.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#MacBook Air M2: 100 GB/s of bandwidth and up to 24 GB
The MacBook Air M2 changes two things compared with the M1 that matter for an LLM: bandwidth rises to 100 GB/s, and memory can reach 24 GB. Apple announces 100 GB/s for the M2, 50% more than the M1; that is the official figure, not a 47% gain as is sometimes reported. Maximum generation speed therefore improves by roughly one-third to one-half depending on the model, and the class of accessible models rises by one level.
- Chip
- Apple M2: 8-core CPU (4 performance, 4 efficiency), 8-core GPU, or 10 cores in the higher-end configuration, 16-core Neural Engine, according to Apple’s specifications.
- Memory
- According to the Apple specifications, M2 MacBook Air models come with 8 GB (configurable to 16 or 24 GB) and 16 GB (configurable to 24 GB). The memory is soldered: you make the choice at purchase.
- Bandwidth
- 100 GB/s, according to Apple, or 50% more than the M1.
- Cooling
- No fan: Apple describes the MacBook Air M2 as a silent, fanless design. Silence comes at the cost of sustained-load performance; see the thermal section.
- Battery
- 52.6 Wh; Apple claims up to 15 hours of wireless web browsing. An LLM that loads the GPU consumes much more than a web page.
#8, 16, or 24 GB: what each tier enables
On an M2 MacBook Air, every gigabyte counts. The Mac guide helps you choose a model that really fits (ch. 4), calculate your memory budget (ch. 5), and measure the impact on heat and battery life (ch. 6).
- Lifetime online access
- PDF + files
- Lifetime updates
A model cannot use all the memory: macOS, your browser, and your applications use some of it. Allow roughly 4 to 5 GB of headroom on 8 GB, 10 to 11 GB on 16 GB, and 16 to 18 GB on 24 GB for the model and its context. These are cautious estimates; compare them with the memory pressure shown in Activity Monitor. A model's size follows the site's guidelines: 3B is about 2 GB, 7-8B about 5 GB, 14B about 9 GB, and 32B about 19 to 20 GB in Q4.
| Memory | Headroom for the model and context | Comfortable tier | On the limit |
|---|---|---|---|
| 8 GB | ≈ 4 to 5 GB | 3-4B | 8B with a very short context |
| 16 GB | ≈ 10 to 11 GB | 8-9B, medium context | 12B with short context |
| 24 GB | ≈ 16 to 18 GB | Comfortable with 12–14B, 8–9B in Q8 or with a long context | 24B in Q4 (≈ 14 GB) with reduced context |
The 24 GB tier is the one that changes the nature of the machine: it opens the door to 24-billion-parameter models in Q4 (about 14 GB of weights), which no M1 Air can handle. It is rare, however: because the memory cannot be upgraded, most M2 Airs in circulation have 8 or 16 GB. Check the Apple menu, “About This Mac,” to see what the machine actually contains before choosing a model.
#The maximum speed the M2 can reach
Generation reads the active weights for every token: maximum speed is bandwidth divided by the weight size in gigabytes. The following table applies the formula to the M2 (100 GB/s) and the M1 (about 68 GB/s). These are calculated ceilings, never measurements: actual speed is lower and depends on the engine, quantization, and context length. They let you check whether a figure reported elsewhere is physically possible.
| Model | Weights | Air M1 (≈ 68 GB/s) | Air M2 (100 GB/s) |
|---|---|---|---|
| 3B | ≈ 2 GB | ≈ 34 tok/s | ≈ 50 tok/s |
| 4B | ≈ 2.5 to 3 GB | ≈ 23 to 27 tok/s | ≈ 33 to 40 tok/s |
| 8B | ≈ 5 GB | ≈ 13 tok/s | ≈ 20 tok/s |
| 12B | ≈ 7.5 GB | ≈ 9 tok/s | ≈ 13 tok/s |
| 24B | ≈ 14 GB | not applicable (memory) | ≈ 7 tok/s |
Two conclusions stand out. First, an 8B model in Q4 on M2 is theoretically below 20 tokens per second: a displayed throughput of 25 or 28 tokens per second for this model on this chip cannot come from standard generation. Second, a 24B model tops out at around 7 tokens per second: it is usable for a measured response, not a fast exchange. For real measurements, look at /benchmarks and third-party benchmarks, checking the chip, quantization, and engine.
#Install a model and check the GPU
- 01Install OllamaDownload it from ollama.com or with Homebrew. The site’s macOS installation guide details the startup options.
- 02Choose the model based on memory4B on 8 GB, 8-9B on 16 GB, 12-14B or a 24B in Q4 on 24 GB. Launch it with ollama run followed by the model name.
- 03Check the GPUIn another terminal, ollama ps displays the CPU/GPU split in the PROCESSOR column. Everything on the GPU is the expected state.
- 04Adjust the contextOllama sets the default context to 4,000 tokens with less than 24 GiB of memory. Increase it gradually rather than all at once: the KV cache uses memory for every token.
Models are installed in ~/.ollama/models. On an M2 Air with a 256 GB SSD, a few 5 to 15 GB models are enough to fill the available space: delete those you no longer use with ollama rm.
A useful rule of thumb for choosing: start with the smallest model that answers your task correctly, then move up one size only if the answers remain insufficient. On a machine with limited bandwidth, each additional billion parameters comes at the cost of speed. Test two candidates on five of your own real queries before deciding: it takes ten minutes and prevents you from downloading a model that is too large.
#24 GB and RAG: what the tier really enables
An assistant for your documents requires three building blocks: an embedding model that indexes, a generation model that answers, and optionally a reranker that sorts the passages. On 8 GB, the combination is tight. On 16 GB, a lightweight embedder and an 8–9B model can coexist with a short context. On 24 GB, you can add a reranker and keep a context of several thousand tokens without writing to the SSD. This is the use case where the next tier really makes sense.
- Before indexing
- Make sure the embedding model fits alongside the generation model: both remain loaded during the request.
- During the request
- Inject a few carefully chosen passages instead of an entire document: context consumes memory and compute time.
- After
- If memory pressure reaches yellow, reduce the number of passes or switch to a smaller model.
#Quantization and KV cache on M2
On a machine limited by memory bandwidth, Q4_K_M remains the default setting: the model is smaller, so it is faster to read. Q5_K_M makes sense when memory allows it and quality comes first. Q8_0 nearly doubles the model size and roughly halves the speed ceiling. For the KV cache, Ollama uses flash attention automatically when the engine and hardware allow it; the OLLAMA_KV_CACHE_TYPE variable quantizes the cache, with f16 as the default. This makes long contexts more accessible on 16 GB, at the cost of a slight loss in precision.
#Battery life, quiet operation, and heat
The Air M2 has no fan: it stays silent, making it suitable for a quiet office or an assistant left running overnight. The chassis dissipates heat passively, so a load lasting several dozen minutes causes the chip to throttle. This page gives neither a duration nor a threshold temperature: no primary source establishes them, and they vary with the room. On battery power, runtime drops much faster than during web browsing because the GPU works continuously; for a long session, plug it in.
- Long-running tasks
- Plug it in, elevate the machine, and avoid applications that load the processor at the same time.
- Power-saving mode
- macOS offers it in the battery settings. It throttles performance: reserve it for travel when battery life matters more than speed.
- Sustained load
- For a local server that responds all day, a Mac mini, which has a fan, is better suited than an Air.
#Stay with M2 or move to an M3, an M4?
The M3 retains the same 100 GB/s bandwidth and the same 24 GB limit on the Air, according to the Apple specifications: for an LLM, moving from M2 to M3 brings almost nothing measurable on the generation side. The M4 increases bandwidth to 120 GB/s and memory to 32 GB, changing both the maximum speed and the model class.
| Criterion | M2 Air | Air M3 | Air M4 |
|---|---|---|---|
| Bandwidth | 100 GB/s | 100 GB/s | 120 GB/s |
| Maximum memory | 24 GB | 24 GB | 32 GB |
| Decision | Keep if 16 GB or 24 GB | Little reason to upgrade from M2 | Switch if you want more than 24 GB |
- Keep your 16 or 24 GB M2
- It handles 8–9B and 12–14B effortlessly; the M4 improves maximum speed by only about 20%.
- Switch if you have 8 GB
- The benefit is in the model class, not the speed: an Air with at least 16 GB lets you move from 4B to 8-9B.
- Switch if you want more than 24 GB
- Only the M4 offers 32 GB on the Air; beyond that, look at the MacBook Pro M4 Pro or Max.
- MacBook Air M1: the limits at 8 and 16 GB
- MacBook Air M3: the same 100 GB/s
- MacBook Air M4: 32 GB and 120 GB/s
- Install Ollama on macOS
- MacBook Pro M4 Pro and Max, up to 128 GB
- Site VRAM calculator
- Source: Apple MacBook Air M2 technical specifications
- Source: Apple, announcement of the M2 chip
- Source: Apple technical specifications for the MacBook Air M4
- Source: Ollama documentation, context
#Frequently asked questions
Can an 8 GB MacBook Air M2 run an 8B model?+
What’s the real difference between the M2 and M3 MacBook Air for an LLM?+
Is the 24 GB MacBook Air M2 worth it for local AI?+
How fast does an 8B LLM run on a MacBook Air M2?+
Does the Air M2 fan turn on with a large model?+
Ollama or LM Studio on an M2 Air?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.