Beginner 11 minMacBook Air

Which LLM on a MacBook Air M2 (8 / 16 / 24 GB) ?

Direct response

The MacBook Air M2 runs 3- to 4-billion-parameter models on 8 GB, 8- to 9-billion-parameter models on 16 GB, and up to a 24B model in Q4 on 24 GB. Its 100 GB/s bandwidth, 50% more than the M1 according to Apple, caps an 8B model in Q4 at about 20 tokens per second; memory determines what can load.

The M2 is the first Air chip to offer 24 GB and 100 GB/s. This page explains which class of model to choose based on 8, 16, or 24 GB, calculates the maximum physically possible speed, and indicates when to move to an M4.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#MacBook Air M2: 100 GB/s of bandwidth and up to 24 GB

The MacBook Air M2 changes two things compared with the M1 that matter for an LLM: bandwidth rises to 100 GB/s, and memory can reach 24 GB. Apple announces 100 GB/s for the M2, 50% more than the M1; that is the official figure, not a 47% gain as is sometimes reported. Maximum generation speed therefore improves by roughly one-third to one-half depending on the model, and the class of accessible models rises by one level.

Chip
Apple M2: 8-core CPU (4 performance, 4 efficiency), 8-core GPU, or 10 cores in the higher-end configuration, 16-core Neural Engine, according to Apple’s specifications.
Memory
According to the Apple specifications, M2 MacBook Air models come with 8 GB (configurable to 16 or 24 GB) and 16 GB (configurable to 24 GB). The memory is soldered: you make the choice at purchase.
Bandwidth
100 GB/s, according to Apple, or 50% more than the M1.
Cooling
No fan: Apple describes the MacBook Air M2 as a silent, fanless design. Silence comes at the cost of sustained-load performance; see the thermal section.
Battery
52.6 Wh; Apple claims up to 15 hours of wireless web browsing. An LLM that loads the GPU consumes much more than a web page.

#8, 16, or 24 GB: what each tier enables

The Mac Kit

On an M2 MacBook Air, every gigabyte counts. The Mac guide helps you choose a model that really fits (ch. 4), calculate your memory budget (ch. 5), and measure the impact on heat and battery life (ch. 6).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A model cannot use all the memory: macOS, your browser, and your applications use some of it. Allow roughly 4 to 5 GB of headroom on 8 GB, 10 to 11 GB on 16 GB, and 16 to 18 GB on 24 GB for the model and its context. These are cautious estimates; compare them with the memory pressure shown in Activity Monitor. A model's size follows the site's guidelines: 3B is about 2 GB, 7-8B about 5 GB, 14B about 9 GB, and 32B about 19 to 20 GB in Q4.

M2 Air: realistic model class by memory (Q4_K_M)
MemoryHeadroom for the model and contextComfortable tierOn the limit
8 GB≈ 4 to 5 GB3-4B8B with a very short context
16 GB≈ 10 to 11 GB8-9B, medium context12B with short context
24 GB≈ 16 to 18 GBComfortable with 12–14B, 8–9B in Q8 or with a long context24B in Q4 (≈ 14 GB) with reduced context

The 24 GB tier is the one that changes the nature of the machine: it opens the door to 24-billion-parameter models in Q4 (about 14 GB of weights), which no M1 Air can handle. It is rare, however: because the memory cannot be upgraded, most M2 Airs in circulation have 8 or 16 GB. Check the Apple menu, “About This Mac,” to see what the machine actually contains before choosing a model.

→
16 GB: the tier that covers most needs
For summarization, writing, coding assistance, and light RAG, a model with 8 to 9 billion parameters in Q4 is sufficient for most common uses. The 24 GB is mainly justified if you want a 24B, a very long context, or an embedder and reranker running alongside the generation model.

#The maximum speed the M2 can reach

Generation reads the active weights for every token: maximum speed is bandwidth divided by the weight size in gigabytes. The following table applies the formula to the M2 (100 GB/s) and the M1 (about 68 GB/s). These are calculated ceilings, never measurements: actual speed is lower and depends on the engine, quantization, and context length. They let you check whether a figure reported elsewhere is physically possible.

Theoretical generation ceiling (weights only in Q4): calculation, not measurement
ModelWeightsAir M1 (≈ 68 GB/s)Air M2 (100 GB/s)
3B≈ 2 GB≈ 34 tok/s≈ 50 tok/s
4B≈ 2.5 to 3 GB≈ 23 to 27 tok/s≈ 33 to 40 tok/s
8B≈ 5 GB≈ 13 tok/s≈ 20 tok/s
12B≈ 7.5 GB≈ 9 tok/s≈ 13 tok/s
24B≈ 14 GBnot applicable (memory)≈ 7 tok/s

Two conclusions stand out. First, an 8B model in Q4 on M2 is theoretically below 20 tokens per second: a displayed throughput of 25 or 28 tokens per second for this model on this chip cannot come from standard generation. Second, a 24B model tops out at around 7 tokens per second: it is usable for a measured response, not a fast exchange. For real measurements, look at /benchmarks and third-party benchmarks, checking the chip, quantization, and engine.

#Install a model and check the GPU

  1. 01
    Install Ollama
    Download it from ollama.com or with Homebrew. The site’s macOS installation guide details the startup options.
  2. 02
    Choose the model based on memory
    4B on 8 GB, 8-9B on 16 GB, 12-14B or a 24B in Q4 on 24 GB. Launch it with ollama run followed by the model name.
  3. 03
    Check the GPU
    In another terminal, ollama ps displays the CPU/GPU split in the PROCESSOR column. Everything on the GPU is the expected state.
  4. 04
    Adjust the context
    Ollama sets the default context to 4,000 tokens with less than 24 GiB of memory. Increase it gradually rather than all at once: the KV cache uses memory for every token.
Terminal
brew install --cask ollama
ollama run qwen3.5:9b
ollama ps

Models are installed in ~/.ollama/models. On an M2 Air with a 256 GB SSD, a few 5 to 15 GB models are enough to fill the available space: delete those you no longer use with ollama rm.

A useful rule of thumb for choosing: start with the smallest model that answers your task correctly, then move up one size only if the answers remain insufficient. On a machine with limited bandwidth, each additional billion parameters comes at the cost of speed. Test two candidates on five of your own real queries before deciding: it takes ten minutes and prevents you from downloading a model that is too large.

#24 GB and RAG: what the tier really enables

An assistant for your documents requires three building blocks: an embedding model that indexes, a generation model that answers, and optionally a reranker that sorts the passages. On 8 GB, the combination is tight. On 16 GB, a lightweight embedder and an 8–9B model can coexist with a short context. On 24 GB, you can add a reranker and keep a context of several thousand tokens without writing to the SSD. This is the use case where the next tier really makes sense.

Before indexing
Make sure the embedding model fits alongside the generation model: both remain loaded during the request.
During the request
Inject a few carefully chosen passages instead of an entire document: context consumes memory and compute time.
After
If memory pressure reaches yellow, reduce the number of passes or switch to a smaller model.

#Quantization and KV cache on M2

On a machine limited by memory bandwidth, Q4_K_M remains the default setting: the model is smaller, so it is faster to read. Q5_K_M makes sense when memory allows it and quality comes first. Q8_0 nearly doubles the model size and roughly halves the speed ceiling. For the KV cache, Ollama uses flash attention automatically when the engine and hardware allow it; the OLLAMA_KV_CACHE_TYPE variable quantizes the cache, with f16 as the default. This makes long contexts more accessible on 16 GB, at the cost of a slight loss in precision.

#Battery life, quiet operation, and heat

The Air M2 has no fan: it stays silent, making it suitable for a quiet office or an assistant left running overnight. The chassis dissipates heat passively, so a load lasting several dozen minutes causes the chip to throttle. This page gives neither a duration nor a threshold temperature: no primary source establishes them, and they vary with the room. On battery power, runtime drops much faster than during web browsing because the GPU works continuously; for a long session, plug it in.

Long-running tasks
Plug it in, elevate the machine, and avoid applications that load the processor at the same time.
Power-saving mode
macOS offers it in the battery settings. It throttles performance: reserve it for travel when battery life matters more than speed.
Sustained load
For a local server that responds all day, a Mac mini, which has a fan, is better suited than an Air.

#Stay with M2 or move to an M3, an M4?

The M3 retains the same 100 GB/s bandwidth and the same 24 GB limit on the Air, according to the Apple specifications: for an LLM, moving from M2 to M3 brings almost nothing measurable on the generation side. The M4 increases bandwidth to 120 GB/s and memory to 32 GB, changing both the maximum speed and the model class.

M2, M3, and M4 Air: the essentials for a local LLM, according to the Apple specs
CriterionM2 AirAir M3Air M4
Bandwidth100 GB/s100 GB/s120 GB/s
Maximum memory24 GB24 GB32 GB
DecisionKeep if 16 GB or 24 GBLittle reason to upgrade from M2Switch if you want more than 24 GB
Keep your 16 or 24 GB M2
It handles 8–9B and 12–14B effortlessly; the M4 improves maximum speed by only about 20%.
Switch if you have 8 GB
The benefit is in the model class, not the speed: an Air with at least 16 GB lets you move from 4B to 8-9B.
Switch if you want more than 24 GB
Only the M4 offers 32 GB on the Air; beyond that, look at the MacBook Pro M4 Pro or Max.

#Frequently asked questions

FAQ
Can an 8 GB MacBook Air M2 run an 8B model?+
Technically yes in Q4 (about 5 GB of weights), but the 4 to 5 GB of headroom macOS leaves is exhausted as soon as the context grows: the system then writes to the SSD and everything slows down. On 8 GB, stick to a model with 3 to 4 billion parameters and keep the context short.
What’s the real difference between the M2 and M3 MacBook Air for an LLM?+
Very little for generation: according to the Apple specifications, the Air M3 offers the same 100 GB/s bandwidth and the same maximum of 24 GB of memory. The M3 mainly brings other improvements, with no direct effect on weight-reading speed. If you have an M2, wait for the M4 instead.
Is the 24 GB MacBook Air M2 worth it for local AI?+
Yes if you want a 24-billion-parameter model in Q4 (about 14 GB) or a RAG setup with an embedder and reranker: these uses do not fit in 16 GB. Otherwise, 16 GB already covers models with 8 to 9 billion parameters. Check the machine's actual memory before buying.
How fast does an 8B LLM run on a MacBook Air M2?+
The theoretical ceiling is bandwidth divided by weight: 100 GB/s divided by approximately 5 GB gives a maximum of 20 tokens per second for an 8B model in Q4. Actual speed is lower, depending on the engine and context. See /benchmarks or third-party benchmarks that specify the chip and quantization.
Does the Air M2 fan turn on with a large model?+
No: there isn't one. Apple describes a fanless, silent design. The tradeoff is thermal: under sustained load, the chip reduces its clock speed and generation slows down silently. For workloads lasting several hours, choose a Mac mini or MacBook Pro.
Ollama or LM Studio on an M2 Air?+
Ollama to get started quickly, script, and integrate with other tools; LM Studio for a graphical interface and to explore Hugging Face models. Both can coexist, but they must not load two models at the same time on 8 or 16 GB: memory is the first limiting factor.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.