Beginner 11 minMacBook Air

Which LLM on a MacBook Air M3 (8 / 16 / 24 GB) ?

Direct response

Local AI on a MacBook Air M3: with 16 GB of unified memory, an 8- to 9-billion-parameter model in Q4 runs comfortably; with 8 GB, stick to 3-4B; with 24 GB, a 24B loads, but the 100 GB/s bandwidth makes it slow. The M3 is no faster than the M2 for text generation: published measurements show nearly identical throughput. Do not confuse it with Apple Intelligence, which is a different service.

The MacBook Air M3 is the most widely used Apple laptop for trying local AI. It is silent, fanless, and runs 8- to 9-billion-parameter models without trouble. Here is what each memory configuration allows, what published benchmarks say about speed, what has not changed since the M2, and which settings really matter.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#MacBook Air M3 in 2026: the technical specs that matter for LLMs

According to Apple's technical specifications, the MacBook Air's M3 chip combines an 8-core CPU (4 performance cores and 4 efficiency cores), an 8- or 10-core GPU, a 16-core Neural Engine, and 100 GB/s of memory bandwidth. The unified memory is soldered: Apple lists two entry-level configurations, 8 GB or 16 GB, both expandable to 24 GB at purchase. The model comes in 13- and 15-inch versions with the same chip.

Bandwidth
100 GB/s, identical to the M2's. This, more than the number of GPU cores, is what caps generation speed.
Cooling
Fanless. The chip can dissipate heat only through the chassis; during a long generation, it eventually reduces its clock speed.
Neural Engine
16 cores, but common local AI tools (Ollama, llama.cpp, LM Studio) use the GPU through Metal, not the Neural Engine.
What's new in M3
Dynamic Caching, which allocates the GPU's local memory in real time. We have no measurement of its effect on LLMs and make no promises about it.
i
What the M3 didn't change
The M3 retains the M2's 100 GB/s. The llama.cpp measurement table confirms it: a Llama 7B in Q4_0 generates 21.34 tokens per second on M3 versus 21.91 on M2 (Q8_0: 12.27 versus 12.21). Moving from an M2 Air to an M3 Air therefore doesn't speed up text generation. Bandwidth changes only with the M4, at 120 GB/s.

#Local AI or Apple Intelligence: two different things

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A lot of research on the MacBook Air M3 concerns “AI” in the broad sense. Apple Intelligence is the suite of features integrated into macOS (summaries, writing, Siri); according to Apple, it works on Macs equipped with an M1 chip or later, including the MacBook Air M3. It doesn't let you choose the model, expose an API, or work with your own documents using an open model.

The local AI discussed in this guide consists of downloading an open model (Qwen, Gemma, Granite, Mistral) and running it with Ollama, LM Studio, or MLX. You choose the model, context size, and what stays on the machine. The two can coexist; memory, however, is shared, and a loaded 9-billion-parameter model reduces what remains for macOS and applications.

#How much memory: 8, 16, or 24 GB

Memory is the primary criterion because a model that spills over does not just slow down: it becomes unusable. On Apple Silicon, the GPU cannot use all unified memory; according to discussions in the llama.cpp repository, Metal allocates between two-thirds and three-quarters of it by default. On a 16 GB machine, plan for about 10 to 12 GB for the model and its context.

What each configuration enables (Q4 weights from the QuelLLM catalog)
MemoryUsable for the model (estimated)What fits comfortablyWhat barely fits
8 GB5 to 6 GBQwen 3.5 4B (2.3 GB), Gemma 4 2B (1.2 GB)Granite 4.2 8B (4.6 GB), short context
16 GB10 to 12 GBQwen 3.5 9B (6 GB), Granite 4.2 8B, Gemma 4 12B (7 GB)Models of 14 GB and larger: no
24 GB16 to 18 GBQwen 3.5 9B in Q8, Gemma 4 12B, 14 to 16 GB modelsMistral Small 3.2 24B (14 GB) or Qwen 3.5 27B (16 GB): they load, but slowly

The “usable” column is an estimate based on the two-thirds to three-quarters rule, not an Apple value. 8 GB makes sense only if you know you'll stick to small models; since memory can't be added later, 16 GB is the sensible choice for local AI. To check whether a model and its context fit, the site's memory calculator gives you the total.

#Which model to choose, and what speed to expect

Generation throughput is limited by bandwidth divided by the size of the weights read for each token. With 100 GB/s, a 4.6 GB model is capped at around 21 tokens per second, a 6 GB model at 16, and a 14 GB model at 7. In published measurements for a Llama 7B, the M3 reaches 21.34 tokens per second in Q4_0, or about 80% of the cap. These are theoretical caps, not guarantees.

Theoretical ceiling of 100 GB/s depending on weight size
Model (Q4)WeightsCeiling (100 ÷ weights)Key takeaway
Qwen 3.5 4B2.3 GBabout 43 t/sVery smooth, with short, fast responses
Granite 4.2 8B4.6 GBabout 21 t/sGood compromise for chat and summarization
Qwen 3.5 9B6 GBabout 16 t/sChoosing 16 GB for quality
Gemma 4 12B7 GBapproximately 14 t/sCorrect on 16 GB, with a moderate context
Mistral Small 3.2 24B14 GBapproximately 7 t/sReadable, but no longer interactive use
Qwen 3.5 27B16 GBabout 6 t/sReserved for long tasks on 24 GB

If the result is too slow, the lever isn't a setting but the model size: going from 9B to 4B roughly doubles throughput. For code, a specialized 7B to 9B model is more useful than a 24B model that is too slow for direct use. The site's speed table (/benchmarks) lists measurements published by hardware.

#Fanless: heat limits long generations

The MacBook Air has no fan. For a short question, you won’t notice. For a summary of several dozen pages or generation lasting several minutes, the chip heats up and eventually reduces its clock speed, and therefore its throughput. The phenomenon depends on the room, workload, and environment; we have no measurement to cite, and it’s best to verify it on your machine with a long generation.

Place the machine on a ventilated surface
An aluminum stand or a hard surface helps the chassis dissipate heat; a couch or bed prevents it.
Plan for pauses
For long tasks, split the work into batches instead of running a single generation that takes several minutes.
Reduce the model
A 4B or 8B model runs significantly cooler than one that continuously saturates the memory bandwidth.
Connect the power
On battery power, macOS may limit performance to save energy.

#Install Ollama and run your first model

Ollama requires macOS Sonoma (14) or later. The simplest option is to download the application from the official website; the installation guide details the Homebrew variant. Once launched, the application listens on port 11434 on the machine, and models can be downloaded with one command.

  1. 01
    Install Ollama
    Download the application or follow the installation guide on macOS, then launch it for the first time.
  2. 02
    Choose a model based on memory
    8 GB: a model with 3 to 4 billion parameters. 16 GB: an 8B (5.2 GB in the Ollama library). 24 GB: a 14B (9.3 GB) or larger.
  3. 03
    Start generation
    Run ollama run followed by the model name: the download happens on the first call.
  4. 04
    Check memory during use
    Open Activity Monitor and select the Memory tab: pressure should remain green. If it turns yellow or red, the model is too large or the context is too long.
First model on 16 GB
ollama run qwen3:8b

The Ollama library lists 5.2 GB for qwen3:8b and 9.3 GB for qwen3:14b. The latter works on a 24 GB MacBook Air, but not on a 16 GB model, where it leaves too little headroom. If you want to test MLX instead of Ollama, the dedicated guide explains the differences and limitations of each tool.

#Two Ollama settings that matter on 16 GB

The first is context size. The official FAQ for Ollama specifies a default window of 4,096 tokens, which can be changed; increasing it consumes memory for the cache. For long documents, increase it gradually instead of setting a maximum value. The second is KV-cache quantization: according to the same FAQ, Ollama accepts the OLLAMA_KV_CACHE_TYPE variable with the value q8_0, which uses about half the memory of the default f16 format, provided Flash Attention is enabled.

Reduce the context cache by half
launchctl setenv OLLAMA_FLASH_ATTENTION 1
launchctl setenv OLLAMA_KV_CACHE_TYPE q8_0
# puis quitter et relancer l'application Ollama

On Mac, the FAQ specifies that when Ollama runs as an application, the variables are set with launchctl and the application must then be restarted. The Mac Apple Silicon optimization guide completes these settings on macOS.

#MLX or Ollama on a MacBook Air M3

MLX is Apple's machine learning framework for its chips. It suits developers who want to control the model from Python or run a local LoRA fine-tune. Ollama, meanwhile, handles downloading, loading, and the API; it is the simplest entry point. The speed differences between the two depend on the model and version: rather than promise a percentage, refer to the site's detailed comparison.

#Should you upgrade to a MacBook Air M4 or later?

The M4 raises bandwidth to 120 GB/s according to the llama.cpp table, 20% more than the M3: in that table, a Llama 7B in Q4_0 goes from 21.34 to 24.11 tokens per second. It's noticeable on 8B models and larger, but not transformative. The M3's real limit remains memory, not speed.

When to keep your M3 Air, when to switch
Your situationRecommendation
M3 Air 16 GB, chat and summarization with 8–9B modelsKeep it: an M4 delivers about 20% higher throughput
8 GB M3 Air, interested in 8B models and largerSwitch to 16 GB or more: memory is the limit
Need a comfortable 24B to 32B modelA Mac with more bandwidth and memory (MacBook Pro, Mac mini M4 Pro)
Long continuous generationsA Mac with a fan, to prevent frequency throttling

#Frequently asked questions

FAQ
Can the MacBook Air M3 run a 70B LLM?+
No, not in any usable way. A 70B model in Q4 weighs about 40 GB, far more than the MacBook Air M3’s maximum 24 GB. Even with extremely aggressive quantization, it would not fit. The reasonable ceiling on this machine is a 24B to 27B model with 24 GB, running slowly, and an 8B–9B model with 16 GB for smooth use.
What speed can you expect on a MacBook Air M3?+
For a Llama 7B in Q4_0, the llama.cpp benchmark table reports 21.34 tokens per second for an M3 with 10 GPU cores. A larger model will be slower; a 4B model will be about twice as fast. These are user measurements on an older model and should be treated as an order of magnitude.
MacBook Air M3 8 GB or MacBook Air M2 16 GB for local AI?+
The M2 16 GB, without hesitation. Both chips have the same 100 GB/s bandwidth, so the M3 does not generate faster, while 16 GB can load an 8–9B model that 8 GB cannot accommodate comfortably. Memory is the deciding factor.
Can you use the Neural Engine to run LLMs?+
Not with common tools. Ollama, llama.cpp, and LM Studio run the model on the GPU through Metal. The 16-core Neural Engine is mainly used for Core ML tasks, such as vision or speech recognition, and is not the way to accelerate a chat LLM on this Mac.
Does the MacBook Air M3 get too hot for local AI?+
It has no fan, so it eventually reduces its clock speed during long generations. You won't notice it with short questions. For long tasks, place it on a heat-dissipating stand, plug it in, and choose a smaller model. A machine with a fan is preferable for continuous use.
Which model should you use for coding on a MacBook Air M3 with 16 GB?+
A 7- to 9-billion-parameter code model in Q4, about 5 to 6 GB, leaves room for your file context. A 14B in Q4 (about 9 GB) is feasible if you accept lower throughput. Beyond that, a long context consumes the remaining headroom and the speed becomes too low for assisted typing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.