Which LLM on a MacBook Air M3 (8 / 16 / 24 GB) ?
Local AI on a MacBook Air M3: with 16 GB of unified memory, an 8- to 9-billion-parameter model in Q4 runs comfortably; with 8 GB, stick to 3-4B; with 24 GB, a 24B loads, but the 100 GB/s bandwidth makes it slow. The M3 is no faster than the M2 for text generation: published measurements show nearly identical throughput. Do not confuse it with Apple Intelligence, which is a different service.
The MacBook Air M3 is the most widely used Apple laptop for trying local AI. It is silent, fanless, and runs 8- to 9-billion-parameter models without trouble. Here is what each memory configuration allows, what published benchmarks say about speed, what has not changed since the M2, and which settings really matter.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#MacBook Air M3 in 2026: the technical specs that matter for LLMs
According to Apple's technical specifications, the MacBook Air's M3 chip combines an 8-core CPU (4 performance cores and 4 efficiency cores), an 8- or 10-core GPU, a 16-core Neural Engine, and 100 GB/s of memory bandwidth. The unified memory is soldered: Apple lists two entry-level configurations, 8 GB or 16 GB, both expandable to 24 GB at purchase. The model comes in 13- and 15-inch versions with the same chip.
- Bandwidth
- 100 GB/s, identical to the M2's. This, more than the number of GPU cores, is what caps generation speed.
- Cooling
- Fanless. The chip can dissipate heat only through the chassis; during a long generation, it eventually reduces its clock speed.
- Neural Engine
- 16 cores, but common local AI tools (Ollama, llama.cpp, LM Studio) use the GPU through Metal, not the Neural Engine.
- What's new in M3
- Dynamic Caching, which allocates the GPU's local memory in real time. We have no measurement of its effect on LLMs and make no promises about it.
#Local AI or Apple Intelligence: two different things
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
A lot of research on the MacBook Air M3 concerns “AI” in the broad sense. Apple Intelligence is the suite of features integrated into macOS (summaries, writing, Siri); according to Apple, it works on Macs equipped with an M1 chip or later, including the MacBook Air M3. It doesn't let you choose the model, expose an API, or work with your own documents using an open model.
The local AI discussed in this guide consists of downloading an open model (Qwen, Gemma, Granite, Mistral) and running it with Ollama, LM Studio, or MLX. You choose the model, context size, and what stays on the machine. The two can coexist; memory, however, is shared, and a loaded 9-billion-parameter model reduces what remains for macOS and applications.
#How much memory: 8, 16, or 24 GB
Memory is the primary criterion because a model that spills over does not just slow down: it becomes unusable. On Apple Silicon, the GPU cannot use all unified memory; according to discussions in the llama.cpp repository, Metal allocates between two-thirds and three-quarters of it by default. On a 16 GB machine, plan for about 10 to 12 GB for the model and its context.
| Memory | Usable for the model (estimated) | What fits comfortably | What barely fits |
|---|---|---|---|
| 8 GB | 5 to 6 GB | Qwen 3.5 4B (2.3 GB), Gemma 4 2B (1.2 GB) | Granite 4.2 8B (4.6 GB), short context |
| 16 GB | 10 to 12 GB | Qwen 3.5 9B (6 GB), Granite 4.2 8B, Gemma 4 12B (7 GB) | Models of 14 GB and larger: no |
| 24 GB | 16 to 18 GB | Qwen 3.5 9B in Q8, Gemma 4 12B, 14 to 16 GB models | Mistral Small 3.2 24B (14 GB) or Qwen 3.5 27B (16 GB): they load, but slowly |
The “usable” column is an estimate based on the two-thirds to three-quarters rule, not an Apple value. 8 GB makes sense only if you know you'll stick to small models; since memory can't be added later, 16 GB is the sensible choice for local AI. To check whether a model and its context fit, the site's memory calculator gives you the total.
#Which model to choose, and what speed to expect
Generation throughput is limited by bandwidth divided by the size of the weights read for each token. With 100 GB/s, a 4.6 GB model is capped at around 21 tokens per second, a 6 GB model at 16, and a 14 GB model at 7. In published measurements for a Llama 7B, the M3 reaches 21.34 tokens per second in Q4_0, or about 80% of the cap. These are theoretical caps, not guarantees.
| Model (Q4) | Weights | Ceiling (100 ÷ weights) | Key takeaway |
|---|---|---|---|
| Qwen 3.5 4B | 2.3 GB | about 43 t/s | Very smooth, with short, fast responses |
| Granite 4.2 8B | 4.6 GB | about 21 t/s | Good compromise for chat and summarization |
| Qwen 3.5 9B | 6 GB | about 16 t/s | Choosing 16 GB for quality |
| Gemma 4 12B | 7 GB | approximately 14 t/s | Correct on 16 GB, with a moderate context |
| Mistral Small 3.2 24B | 14 GB | approximately 7 t/s | Readable, but no longer interactive use |
| Qwen 3.5 27B | 16 GB | about 6 t/s | Reserved for long tasks on 24 GB |
If the result is too slow, the lever isn't a setting but the model size: going from 9B to 4B roughly doubles throughput. For code, a specialized 7B to 9B model is more useful than a 24B model that is too slow for direct use. The site's speed table (/benchmarks) lists measurements published by hardware.
#Fanless: heat limits long generations
The MacBook Air has no fan. For a short question, you won’t notice. For a summary of several dozen pages or generation lasting several minutes, the chip heats up and eventually reduces its clock speed, and therefore its throughput. The phenomenon depends on the room, workload, and environment; we have no measurement to cite, and it’s best to verify it on your machine with a long generation.
- Place the machine on a ventilated surface
- An aluminum stand or a hard surface helps the chassis dissipate heat; a couch or bed prevents it.
- Plan for pauses
- For long tasks, split the work into batches instead of running a single generation that takes several minutes.
- Reduce the model
- A 4B or 8B model runs significantly cooler than one that continuously saturates the memory bandwidth.
- Connect the power
- On battery power, macOS may limit performance to save energy.
#Install Ollama and run your first model
Ollama requires macOS Sonoma (14) or later. The simplest option is to download the application from the official website; the installation guide details the Homebrew variant. Once launched, the application listens on port 11434 on the machine, and models can be downloaded with one command.
- 01Install OllamaDownload the application or follow the installation guide on macOS, then launch it for the first time.
- 02Choose a model based on memory8 GB: a model with 3 to 4 billion parameters. 16 GB: an 8B (5.2 GB in the Ollama library). 24 GB: a 14B (9.3 GB) or larger.
- 03Start generationRun ollama run followed by the model name: the download happens on the first call.
- 04Check memory during useOpen Activity Monitor and select the Memory tab: pressure should remain green. If it turns yellow or red, the model is too large or the context is too long.
The Ollama library lists 5.2 GB for qwen3:8b and 9.3 GB for qwen3:14b. The latter works on a 24 GB MacBook Air, but not on a 16 GB model, where it leaves too little headroom. If you want to test MLX instead of Ollama, the dedicated guide explains the differences and limitations of each tool.
#Two Ollama settings that matter on 16 GB
The first is context size. The official FAQ for Ollama specifies a default window of 4,096 tokens, which can be changed; increasing it consumes memory for the cache. For long documents, increase it gradually instead of setting a maximum value. The second is KV-cache quantization: according to the same FAQ, Ollama accepts the OLLAMA_KV_CACHE_TYPE variable with the value q8_0, which uses about half the memory of the default f16 format, provided Flash Attention is enabled.
On Mac, the FAQ specifies that when Ollama runs as an application, the variables are set with launchctl and the application must then be restarted. The Mac Apple Silicon optimization guide completes these settings on macOS.
#MLX or Ollama on a MacBook Air M3
MLX is Apple's machine learning framework for its chips. It suits developers who want to control the model from Python or run a local LoRA fine-tune. Ollama, meanwhile, handles downloading, loading, and the API; it is the simplest entry point. The speed differences between the two depend on the model and version: rather than promise a percentage, refer to the site's detailed comparison.
#Should you upgrade to a MacBook Air M4 or later?
The M4 raises bandwidth to 120 GB/s according to the llama.cpp table, 20% more than the M3: in that table, a Llama 7B in Q4_0 goes from 21.34 to 24.11 tokens per second. It's noticeable on 8B models and larger, but not transformative. The M3's real limit remains memory, not speed.
| Your situation | Recommendation |
|---|---|
| M3 Air 16 GB, chat and summarization with 8–9B models | Keep it: an M4 delivers about 20% higher throughput |
| 8 GB M3 Air, interested in 8B models and larger | Switch to 16 GB or more: memory is the limit |
| Need a comfortable 24B to 32B model | A Mac with more bandwidth and memory (MacBook Pro, Mac mini M4 Pro) |
| Long continuous generations | A Mac with a fan, to prevent frequency throttling |
- MacBook Air M4: the successor
- MacBook Air M2: the previous generation
- MacBook Pro M2: with fan
- macOS settings for Apple Silicon
- Source: MacBook Air M3 technical specifications (Apple)
- Source: llama.cpp measurements on Apple chips
- Source: official Ollama FAQ
#Frequently asked questions
Can the MacBook Air M3 run a 70B LLM?+
What speed can you expect on a MacBook Air M3?+
MacBook Air M3 8 GB or MacBook Air M2 16 GB for local AI?+
Can you use the Neural Engine to run LLMs?+
Does the MacBook Air M3 get too hot for local AI?+
Which model should you use for coding on a MacBook Air M3 with 16 GB?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.