Which LLM on a MacBook Air M1 (8 / 16 GB) ?
The M1 MacBook Air runs models with 3 to 4 billion parameters on 8 GB, and models with 8 to 9 billion on 16 GB, using Q4_K_M quantization. The limitation is not compute power but unified memory (8 or 16 GB, not expandable) and bandwidth of about 68 GB/s, which caps an 8B model at around a dozen tokens per second.
The MacBook Air M1 remains an honest entry point into local AI, provided you choose the right model class. This page decides for you based on your memory, calculates the maximum speed the machine can physically reach, and flags when it's time to move on.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#MacBook Air M1: what the machine really offers for an LLM
For a local LLM, the MacBook Air M1 comes down to three figures: unified memory (8 GB standard, 16 GB optional, never expandable after purchase), memory bandwidth (around 68 GB/s), and the lack of a fan. Memory determines what can be loaded, bandwidth determines generation speed, and passive cooling determines how long that speed lasts.
- Chip
- Apple M1 (2020): 8 CPU cores (4 performance, 4 efficiency), 7 or 8 GPU cores depending on the configuration, and a 16-core Neural Engine, according to Apple's technical specifications.
- Memory
- 8 GB of unified memory, configurable to 16 GB. The CPU and GPU draw from the same pool: there is no separate VRAM to add.
- Bandwidth
- Apple does not publish it in the M1 specifications. It does, however, announce 100 GB/s for the M2, 50% more than the M1: that puts the M1 at around 67–68 GB/s, the figure commonly cited (LPDDR4X).
- Cooling
- No fan: Apple presents the MacBook Air M1 as completely silent, regardless of the task. The downside is thermal; see below.
- Battery
- 49.9 Wh battery; sustained inference drains it much faster than the web browsing on which Apple bases its advertised 15 hours.
Apple has not been sold new by the manufacturer for a long time; you can find it used, and prices vary too much to list here. The site's hardware pages (/materiel-ia) provide current reference points if you're choosing between a used Air and a newer machine. The 2020 Mac mini contains the same M1 chip and has the same memory and bandwidth limits; it adds a fan, so sustained speeds are more stable.
#8 GB or 16 GB: what the RAM enables and prevents
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
With 8 GB, the system, browser, and applications already share the same space as the model. That leaves roughly 4 to 5 GB for an LLM, including context: this is a rule-of-thumb estimate, not a measurement, since macOS adjusts its own consumption. With 16 GB, the margin rises to around 10 to 11 GB, making models with 8 to 9 billion parameters feasible.
| Model | Weights in Q4 | Air M1 8 GB | M1 Air 16 GB |
|---|---|---|---|
| 3B (Granite 4.1 3B, Llama 3.2 3B) | ≈ 2 to 2.5 GB | Comfortable | Comfortable |
| 4B (Qwen 3.5 4B, Phi-4 mini) | ≈ 2.5 to 3 GB | Comfortable, short context | Comfortable, long context |
| 7–9B (Granite 4.1 8B, Qwen 3.5 9B) | ≈ 5 to 6 GB | Barely enough: swap likely | Good compromise |
| 12B (Gemma 4 12B) | ≈ 7 to 8 GB | No | Limitation: short context only |
| 24B and above | ≥ 14 GB | No | No |
The last row's ranking doesn't depend on speed but on a hard boundary: a 14 GB model won't fit in 16 GB once the system is loaded. If you're unsure of a model's exact size, the site's VRAM calculator shows its footprint based on quantization and context.
#Which models to choose, depending on the use case
The M1 should be judged by the model’s size, not its brand. With 8 GB, target a 3B or 4B in Q4_K_M for text summarization, rewriting, question answering over an excerpt, and simple code. With 16 GB, an 8- to 9-billion-parameter model offers a clear gain in accuracy and ability to follow long instructions, at the cost of slower generation. Model names change every month; the key is to remember the size class and check the site’s catalog for the current status.
- Writing, summarization, translation
- A 4B on 8 GB, an 8–9B on 16 GB. French writing quality depends more on the model family than on size: test two candidates on your own texts.
- Code
- A 4B is enough for short functions and explanations; for an entire project, the Air M1's memory requires short contexts, so a local coding assistant remains limited.
- Documents and RAG
- A lightweight embedder plus a 4B (8 GB) or 8–9B (16 GB) generation model. Injected context consumes memory: keep the number of passages small and choose them carefully.
- Vision
- Multimodal models add the image encoder to the model's memory. On 8 GB, reserve them for small sizes; on 16 GB, they run but slowly.
#How many tokens per second: the theoretical ceiling
During generation, the processor rereads roughly all of the model's active weights for each token produced. Maximum speed is therefore capped by bandwidth divided by weight size. At about 68 GB/s on an M1, a model weighing 5 GB cannot exceed roughly 13 tokens per second, regardless of the available GPU cores. This is a calculated ceiling, not a measurement: actual speed is lower because computation, the KV cache, and the system also consume bandwidth.
| Model | Weights in Q4 | Compute | Cap |
|---|---|---|---|
| 3B | ≈ 2 GB | 68 ÷ 2 | ≈ 34 tok/s |
| 4B | ≈ 2.5 to 3 GB | 68 ÷ 2,5 à 3 | ≈ 23 to 27 tok/s |
| 8B | ≈ 5 GB | 68 ÷ 5 | ≈ 13 tok/s |
| 9B | ≈ 6 GB | 68 ÷ 6 | ≈ 11 tok/s |
| 12B | ≈ 7.5 GB | 68 ÷ 7,5 | ≈ 9 tok/s |
This calculation serves as a filter. Any claim of more than 30 tokens per second on an 8-billion-parameter model in Q4 on an M1 is physically incompatible with the bandwidth, except for a model mixing experts where only a fraction is active. For real-world measurements, consult the /benchmarks page on the site and benchmarks published by third parties, checking the chip, quantization, and engine used.
#Install and verify in a few minutes
Ollama is the simplest route: it installs the engine, downloads the models, and detects the Mac’s GPU without configuration. Models are stored in ~/.ollama/models on macOS: allow several gigabytes per model and keep this in mind when using a 256 GB SSD.
- 01Install OllamaDownload the application from ollama.com, or use Homebrew with the command below. The site's macOS installation guide explains automatic startup in detail.
- 02Run a model that fits in RAMOn 8 GB, a 4B; on 16 GB, an 8-9B. The first launch downloads the model; subsequent launches are immediate.
- 03Verify that the GPU is workingIn a second terminal, type ollama ps: the PROCESSOR column should show 100% GPU. A CPU/GPU split indicates that the model exceeds the memory reserved for the GPU.
- 04Set the contextOllama chooses a default context based on available memory; below 24 GB, it is 4,000 tokens. Increase it only if needed: each context token consumes memory.
LM Studio is the alternative with a full graphical interface, useful if the terminal puts you off. It uses slightly more memory than Ollama alone, which matters on 8 GB. MLX, Apple's framework, uses the Mac's shared memory with no copies between CPU and GPU; MLX and llama.cpp settings are covered in the macOS optimization guide.
#Long context: the 8 GB trap
The model’s weight is only part of the memory footprint. Each context token adds an entry to the KV cache, which grows with the conversation length. Ollama automatically uses flash attention when the engine and hardware allow it, limiting this growth. It also lets you quantize the KV cache through the OLLAMA_KV_CACHE_TYPE variable, whose default is f16. On an 8 GB M1 Air, this option can make the difference between a 4,000-token context and an 8,000-token context, at the cost of a slight loss in precision.
The guide to KV cache quantization explains when this option is worthwhile. The practical rule: on 8 GB, keep the context short and inject only the relevant passages; on 16 GB, a context of 8,000 tokens remains reasonable with an 8-9B model.
#The heat generated by a fanless Air
Without a fan, the MacBook Air M1 dissipates heat through its chassis. During brief use (one question, one answer), there is nothing to report. Under sustained use, such as processing a batch for several dozen minutes or indexing documents, the chip reduces its clock speed to protect itself: performance drops without the machine shutting down. This page gives no threshold in minutes or degrees: it depends on the room and the workload, and no reliable source specifies one.
- Raise the machine
- A computer stand improves airflow under the chassis.
- Close resource-intensive applications
- Loaded browsers and video conferencing apps use processor resources in parallel and heat up the same chip.
- Prefer a smaller model over Q8
- Less memory to reread, less energy used per token.
- Plug it in for long-running tasks
- On battery power, macOS may throttle performance further. This depends on the power mode selected in Settings.
#Limits and signals that it's time to move to another machine
The M1 is not ruled out because of its age, but because of its memory. The criteria that should make you look elsewhere are simple and measurable.
- You want a model larger than 12B
- No M1 Air configuration can accommodate it comfortably. Consider the MacBook Air M4 (up to 32 GB) or a MacBook Pro Pro/Max.
- You work with large documents
- Long context saturates memory before compute power. More RAM is the only real remedy.
- You run jobs lasting several hours
- A Mac mini, which has a fan, or a desktop PC handles sustained loads better.
- You’re interested in fine-tuning
- An M1 Air isn't designed for this; prefer an occasional cloud service.
| Criterion | MacBook Air M1 | MacBook Air M2 | MacBook Air M4 |
|---|---|---|---|
| Maximum memory | 16 GB | 24 GB | 32 GB |
| Bandwidth | ≈ 68 GB/s | 100 GB/s | 120 GB/s |
| Comfortable Q4 model | 8-9B | 8–9B, 14B in 24 GB | 14B, 24B in 32 GB |
- MacBook Air M2: 24 GB and 100 GB/s of bandwidth
- MacBook Air M4: up to 32 GB of memory
- Install Ollama on macOS, step by step
- Quantizing the KV cache to save memory
- Which LLM for 8 GB of memory
- Site VRAM calculator
- Source: Apple for the MacBook Air M1
- Source: Apple, M2 chip announcement (bandwidth)
- Source: Ollama documentation, context length
- Source: Ollama FAQ, GPU, KV cache, and flash attention
#Frequently asked questions
Is 8 GB on a MacBook Air M1 enough for a local LLM?+
What speed should you expect from an 8B LLM on a MacBook Air M1?+
Can you run a 70B on a MacBook Air M1?+
Should you choose Ollama, LM Studio, or MLX on M1?+
Does the MacBook Air M1 get very hot with an LLM?+
Is it worth upgrading to an Air M4 for local AI?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.