Beginner 12 minMacBook Air

Which LLM on a MacBook Air M1 (8 / 16 GB) ?

Direct response

The M1 MacBook Air runs models with 3 to 4 billion parameters on 8 GB, and models with 8 to 9 billion on 16 GB, using Q4_K_M quantization. The limitation is not compute power but unified memory (8 or 16 GB, not expandable) and bandwidth of about 68 GB/s, which caps an 8B model at around a dozen tokens per second.

The MacBook Air M1 remains an honest entry point into local AI, provided you choose the right model class. This page decides for you based on your memory, calculates the maximum speed the machine can physically reach, and flags when it's time to move on.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#MacBook Air M1: what the machine really offers for an LLM

For a local LLM, the MacBook Air M1 comes down to three figures: unified memory (8 GB standard, 16 GB optional, never expandable after purchase), memory bandwidth (around 68 GB/s), and the lack of a fan. Memory determines what can be loaded, bandwidth determines generation speed, and passive cooling determines how long that speed lasts.

Chip
Apple M1 (2020): 8 CPU cores (4 performance, 4 efficiency), 7 or 8 GPU cores depending on the configuration, and a 16-core Neural Engine, according to Apple's technical specifications.
Memory
8 GB of unified memory, configurable to 16 GB. The CPU and GPU draw from the same pool: there is no separate VRAM to add.
Bandwidth
Apple does not publish it in the M1 specifications. It does, however, announce 100 GB/s for the M2, 50% more than the M1: that puts the M1 at around 67–68 GB/s, the figure commonly cited (LPDDR4X).
Cooling
No fan: Apple presents the MacBook Air M1 as completely silent, regardless of the task. The downside is thermal; see below.
Battery
49.9 Wh battery; sustained inference drains it much faster than the web browsing on which Apple bases its advertised 15 hours.

Apple has not been sold new by the manufacturer for a long time; you can find it used, and prices vary too much to list here. The site's hardware pages (/materiel-ia) provide current reference points if you're choosing between a used Air and a newer machine. The 2020 Mac mini contains the same M1 chip and has the same memory and bandwidth limits; it adds a fan, so sustained speeds are more stable.

#8 GB or 16 GB: what the RAM enables and prevents

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

With 8 GB, the system, browser, and applications already share the same space as the model. That leaves roughly 4 to 5 GB for an LLM, including context: this is a rule-of-thumb estimate, not a measurement, since macOS adjusts its own consumption. With 16 GB, the margin rises to around 10 to 11 GB, making models with 8 to 9 billion parameters feasible.

What fits on an M1 Air depending on memory (Q4_K_M, weights only according to the site's benchmarks, plus context)
ModelWeights in Q4Air M1 8 GBM1 Air 16 GB
3B (Granite 4.1 3B, Llama 3.2 3B)≈ 2 to 2.5 GBComfortableComfortable
4B (Qwen 3.5 4B, Phi-4 mini)≈ 2.5 to 3 GBComfortable, short contextComfortable, long context
7–9B (Granite 4.1 8B, Qwen 3.5 9B)≈ 5 to 6 GBBarely enough: swap likelyGood compromise
12B (Gemma 4 12B)≈ 7 to 8 GBNoLimitation: short context only
24B and above≥ 14 GBNoNo

The last row's ranking doesn't depend on speed but on a hard boundary: a 14 GB model won't fit in 16 GB once the system is loaded. If you're unsure of a model's exact size, the site's VRAM calculator shows its footprint based on quantization and context.

!
Swap is the real warning sign
When macOS runs out of RAM, it writes to the SSD: generation slows significantly and drive wear increases. Open Activity Monitor and select the Memory tab: if “memory pressure” turns yellow or red during generation, choose a smaller model or reduce the context. That’s more effective than closing tabs one by one.

#Which models to choose, depending on the use case

The M1 should be judged by the model’s size, not its brand. With 8 GB, target a 3B or 4B in Q4_K_M for text summarization, rewriting, question answering over an excerpt, and simple code. With 16 GB, an 8- to 9-billion-parameter model offers a clear gain in accuracy and ability to follow long instructions, at the cost of slower generation. Model names change every month; the key is to remember the size class and check the site’s catalog for the current status.

Writing, summarization, translation
A 4B on 8 GB, an 8–9B on 16 GB. French writing quality depends more on the model family than on size: test two candidates on your own texts.
Code
A 4B is enough for short functions and explanations; for an entire project, the Air M1's memory requires short contexts, so a local coding assistant remains limited.
Documents and RAG
A lightweight embedder plus a 4B (8 GB) or 8–9B (16 GB) generation model. Injected context consumes memory: keep the number of passages small and choose them carefully.
Vision
Multimodal models add the image encoder to the model's memory. On 8 GB, reserve them for small sizes; on 16 GB, they run but slowly.

#How many tokens per second: the theoretical ceiling

During generation, the processor rereads roughly all of the model's active weights for each token produced. Maximum speed is therefore capped by bandwidth divided by weight size. At about 68 GB/s on an M1, a model weighing 5 GB cannot exceed roughly 13 tokens per second, regardless of the available GPU cores. This is a calculated ceiling, not a measurement: actual speed is lower because computation, the KV cache, and the system also consume bandwidth.

Theoretical maximum generation on an M1 Air (≈ 68 GB/s, weights only in Q4): calculation, not measurement
ModelWeights in Q4ComputeCap
3B≈ 2 GB68 ÷ 2≈ 34 tok/s
4B≈ 2.5 to 3 GB68 ÷ 2,5 à 3≈ 23 to 27 tok/s
8B≈ 5 GB68 ÷ 5≈ 13 tok/s
9B≈ 6 GB68 ÷ 6≈ 11 tok/s
12B≈ 7.5 GB68 ÷ 7,5≈ 9 tok/s

This calculation serves as a filter. Any claim of more than 30 tokens per second on an 8-billion-parameter model in Q4 on an M1 is physically incompatible with the bandwidth, except for a model mixing experts where only a fraction is active. For real-world measurements, consult the /benchmarks page on the site and benchmarks published by third parties, checking the chip, quantization, and engine used.

i
Why Q4 rather than Q8 on M1
A Q8 weighs almost twice as much as a Q4: at equal bandwidth, it generates about half as fast and uses more memory. On 8-9B models, the quality loss from Q4_K_M is generally small. The quantization guide explains the choice between Q4, Q5, and Q8.

#Install and verify in a few minutes

Ollama is the simplest route: it installs the engine, downloads the models, and detects the Mac’s GPU without configuration. Models are stored in ~/.ollama/models on macOS: allow several gigabytes per model and keep this in mind when using a 256 GB SSD.

  1. 01
    Install Ollama
    Download the application from ollama.com, or use Homebrew with the command below. The site's macOS installation guide explains automatic startup in detail.
  2. 02
    Run a model that fits in RAM
    On 8 GB, a 4B; on 16 GB, an 8-9B. The first launch downloads the model; subsequent launches are immediate.
  3. 03
    Verify that the GPU is working
    In a second terminal, type ollama ps: the PROCESSOR column should show 100% GPU. A CPU/GPU split indicates that the model exceeds the memory reserved for the GPU.
  4. 04
    Set the context
    Ollama chooses a default context based on available memory; below 24 GB, it is 4,000 tokens. Increase it only if needed: each context token consumes memory.
Terminal
brew install --cask ollama
ollama run qwen3.5:4b
ollama ps

LM Studio is the alternative with a full graphical interface, useful if the terminal puts you off. It uses slightly more memory than Ollama alone, which matters on 8 GB. MLX, Apple's framework, uses the Mac's shared memory with no copies between CPU and GPU; MLX and llama.cpp settings are covered in the macOS optimization guide.

#Long context: the 8 GB trap

The model’s weight is only part of the memory footprint. Each context token adds an entry to the KV cache, which grows with the conversation length. Ollama automatically uses flash attention when the engine and hardware allow it, limiting this growth. It also lets you quantize the KV cache through the OLLAMA_KV_CACHE_TYPE variable, whose default is f16. On an 8 GB M1 Air, this option can make the difference between a 4,000-token context and an 8,000-token context, at the cost of a slight loss in precision.

Terminal
# Quantifier le cache KV en q8_0 avant de démarrer le serveur
export OLLAMA_KV_CACHE_TYPE=q8_0
export OLLAMA_FLASH_ATTENTION=1
ollama serve

The guide to KV cache quantization explains when this option is worthwhile. The practical rule: on 8 GB, keep the context short and inject only the relevant passages; on 16 GB, a context of 8,000 tokens remains reasonable with an 8-9B model.

#The heat generated by a fanless Air

Without a fan, the MacBook Air M1 dissipates heat through its chassis. During brief use (one question, one answer), there is nothing to report. Under sustained use, such as processing a batch for several dozen minutes or indexing documents, the chip reduces its clock speed to protect itself: performance drops without the machine shutting down. This page gives no threshold in minutes or degrees: it depends on the room and the workload, and no reliable source specifies one.

Raise the machine
A computer stand improves airflow under the chassis.
Close resource-intensive applications
Loaded browsers and video conferencing apps use processor resources in parallel and heat up the same chip.
Prefer a smaller model over Q8
Less memory to reread, less energy used per token.
Plug it in for long-running tasks
On battery power, macOS may throttle performance further. This depends on the power mode selected in Settings.

#Limits and signals that it's time to move to another machine

The M1 is not ruled out because of its age, but because of its memory. The criteria that should make you look elsewhere are simple and measurable.

You want a model larger than 12B
No M1 Air configuration can accommodate it comfortably. Consider the MacBook Air M4 (up to 32 GB) or a MacBook Pro Pro/Max.
You work with large documents
Long context saturates memory before compute power. More RAM is the only real remedy.
You run jobs lasting several hours
A Mac mini, which has a fan, or a desktop PC handles sustained loads better.
You’re interested in fine-tuning
An M1 Air isn't designed for this; prefer an occasional cloud service.
M1 Air or later generation: what changes for a local LLM
CriterionMacBook Air M1MacBook Air M2MacBook Air M4
Maximum memory16 GB24 GB32 GB
Bandwidth≈ 68 GB/s100 GB/s120 GB/s
Comfortable Q4 model8-9B8–9B, 14B in 24 GB14B, 24B in 32 GB

#Frequently asked questions

FAQ
Is 8 GB on a MacBook Air M1 enough for a local LLM?+
Yes for 3- to 4-billion-parameter models in Q4, which weigh about 2 to 3 GB and leave the system some headroom. An 8-billion-parameter model in Q4 (about 5 GB) remains possible but causes macOS to write to the SSD as soon as the context grows. For regular use, 16 GB is much more comfortable.
What speed should you expect from an 8B LLM on a MacBook Air M1?+
The theoretical ceiling is calculated by dividing bandwidth (about 68 GB/s) by the model's weight: for an 8B model in Q4 (about 5 GB), that gives roughly 13 tokens per second at most. Actual speed is lower. This calculation is useful for ruling out unrealistic figures, not for predicting your machine's performance.
Can you run a 70B on a MacBook Air M1?+
No. A 70-billion-parameter model weighs about 40 GB in Q4, more than twice the maximum memory of an M1 Air (16 GB). Even with very aggressive quantization, it does not load properly. The realistic size class on this machine tops out at around 8 to 12 billion parameters.
Should you choose Ollama, LM Studio, or MLX on M1?+
Ollama to get started quickly and script, LM Studio if you prefer a full graphical interface, MLX if you want to use the Apple framework directly from Python. On 8 GB, the memory overhead of a graphical interface matters: start with Ollama, then switch only if a specific need arises.
Does the MacBook Air M1 get very hot with an LLM?+
It has no fan, so it heats up and reduces its clock speed under sustained load. For short exchanges, the effect is negligible. For long workloads, elevate the machine, connect the power adapter, and prefer a smaller model. If you run tasks lasting several hours, a Mac mini with a fan is a better fit.
Is it worth upgrading to an Air M4 for local AI?+
Yes, if memory is holding you back: the M4 Air goes up to 32 GB and offers 120 GB/s of bandwidth versus about 68 GB/s. If you stick to 4- to 8-billion-parameter models for text, the 16 GB M1 remains usable; the benefit of upgrading is then mainly convenience.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.