Which LLM on an iMac M4 (16 / 24 / 32 GB) ?
The iMac M4 comfortably runs 8- to 14-billion-parameter models in Q4: its M4 chip provides 120 GB/s of memory bandwidth, and a public benchmark reports 24 tokens per second on a 7B in Q4_0 with the 10-core GPU. Unified memory (16, 24, or 32 GB) sets the maximum size: 16 GB for a 14B, 24 GB for a slow 24B, and 32 GB for an even slower 30B.
An iMac is primarily a display; for local AI, it is mainly an M4 chip and a quantity of unified memory that can no longer be changed. This guide quantifies what each configuration can run, at what speed according to public measurements, and when a Mac mini is the better purchase.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What an iMac M4 can do for local AI
The iMac M4 runs 8- to 14-billion-parameter models well and accepts larger models at reduced speed. Three figures summarize the essentials. The M4 chip delivers 120 GB/s of memory bandwidth according to the Apple specifications, implying a theoretical ceiling of about 31 tokens per second on a 7B in Q4_0. The public llama.cpp repository benchmark on an M4 with a 10-core GPU reports 24.11 tokens per second during generation, or 77% of that ceiling. Finally, unified memory ranges from 16 to 32 GB and is chosen at purchase: it alone determines the possible model size. What this iMac lacks is bandwidth, not capacity.
- Chip
- Apple M4; the Apple specification does not offer a Pro or Max variant for the iMac.
- Bandwidth
- 120 GB/s, regardless of the configuration.
- Unified memory
- 16 GB standard; 24 GB optional on all versions, 32 GB on the 4-port model.
- Screen
- 24-inch 4.5K (4480 × 2520), which leaves little reason to use a larger GPU: memory is the limit.
- Pricing
- Not indicated here: Apple modifies them; see the site's hardware page.
#Real-world configurations: two models, not three
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
The Apple page distinguishes two families. The 2-port model has an 8-core CPU and an 8-core GPU, with 16 GB of memory expandable to 24 GB. The 4-port model has a 10-core CPU and a 10-core GPU, with 16 GB expandable to 24 or 32 GB. The 120 GB/s bandwidth is identical. A common error, present in the older version of this page, is to assume that an 8 GB configuration still exists: the base memory is 16 GB. For an LLM, the key point is therefore to choose the 4-port model if you want 32 GB.
| Version | GPU | Unified memory | Bandwidth |
|---|---|---|---|
| 2 ports | 8 cores | 16 GB, 24 GB option | 120 GB/s |
| 4 ports | 10 cores | 16 GB, with 24 and 32 GB options | 120 GB/s |
For text generation, the difference between 8 and 10 GPU cores is probably small: speed depends on bandwidth, which is identical. No public measurements from the reference table exist for the 8-core variant, so this claim remains a hypothesis. If you're unsure, pay for memory rather than cores.
#16, 24, or 32 GB: the memory budget
On a Mac, memory is shared between the system and the GPU, and macOS gives the GPU only a fraction of it. The iogpu.wired_limit_mb setting defaults to 0, meaning the kernel derives the limit from the installed memory. On a 32 GB Mac measured by ModelPiper, Metal made 24.96 GiB available to the GPU, or 78% of memory. 75% is commonly cited. For an iMac, therefore, expect roughly 11 to 12 GB for a 16 GB model, 17 to 18 GB for 24 GB, and 24 to 25 GB for 32 GB: these are estimates, so verify them on your machine with the sysctl iogpu.wired_limit_mb command and the Metal tool described by ModelPiper.
| Configuration | Memory available to the GPU | Largest reasonable model |
|---|---|---|
| 16 GB | ≈ 11 to 12 GB | A 14B in Q4 (≈ 9 GB), medium context |
| 24 GB | ≈ 17 to 18 GB | A 24B in Q4 (≈ 14 GB), short context |
| 32 GB | ≈ 24 to 25 GB | A 30B in Q4 (≈ 17 to 18 GB), medium context |
To go beyond these limits, you can increase the share allocated to the GPU. The guide to optimizing Macs with Apple Silicon details this setting and its risks; we do not repeat it here.
- Get the most out of a Mac Apple Silicon: macOS settings
- Memory calculator for a given model and context
#Which models for which memory capacity
The Q4 weights below come from the QuelLLM catalog. They assume the model fits in the usable memory shown in the previous table, including the context. The “smoothness” column applies the 120 GB/s ceiling to each model: these are calculated estimates, not measurements.
| Model | Weights | Minimum requirements | Expected smoothness on M4 |
|---|---|---|---|
| Llama 3.1 8B | ≈ 6 GB | 16 GB | Smooth (about 15 tokens/s) |
| Gemma 4 12B | ≈ 7 GB | 16 GB | Fair (around 13 tokens/s) |
| Phi-4 14B | ≈ 9 GB | 16 GB | Decent (about 10 tokens/s) |
| gpt-oss 20B | ≈ 13 GB | 24 GB | Variable (MoE: faster than a dense 20B model) |
| Devstral Small 2 24B | ≈ 14 GB | 24 GB | Slow (approximately 6 to 7 tokens/s) |
| Gemma 4 31B | ≈ 18 GB | 32 GB | Very slow (about 5 tokens/s) |
The key is not to confuse “fits” with “is comfortable.” A dense 24B model on an M4 base remains readable but slow, at a few words per second, whereas a mixture-of-experts (MoE) model such as gpt-oss 20B or Gemma 4 26B-A4B reads only a few billion parameters at each token, so it outpaces a comparably sized dense model. For a 24 or 32 GB iMac, an MoE is often the best choice.
A common trade-off on 24 GB is a 12B in Q8 (about 13 GB) or a 24B in Q4 (about 14 GB). Both use almost the same amount of memory, so they generate at roughly the same speed, around 7 tokens per second based on the bandwidth ceiling. The latter offers greater reasoning capacity, while the former stays closer to the original model. With equal bandwidth, model size in gigabytes determines throughput, not the number of parameters.
#Speed: bandwidth is the deciding factor
The only public reference measurement is the llama.cpp repository table on Apple Silicon, which tests a 7B LLaMA in Q4_0. For an M4 with a 10-core GPU and 120 GB/s, it reports 221,29 tokens per second for prompt processing and 24,11 for generation. An M4 Pro at 273 GB/s with a 16-core GPU reports 49,64 for generation, a little more than twice as much. This ratio is close to that of the memory bandwidths (2,3 times): a sign that generation depends on memory, not on the number of GPU cores.
| Chip | GPU | Bandwidth | Generation tg (t/s) |
|---|---|---|---|
| M4 (4-port iMac) | 10 cores | 120 GB/s | 24,11 |
| M4 Pro (Mac mini M4 Pro) | 16 cores | 273 GB/s | 49,64 |
These measurements date back to an early version of llama.cpp; recent versions and MLX may do better. Treat them as an order of magnitude. The previous version of this page reported 28 to 31 tokens per second for a 9B in Q5: that figure would exceed the 120 GB/s ceiling for a model of this size, so we removed it.
#Three realistic use cases based on memory
A memory budget is easier to understand when applied to a use case. In each case, add the model weights, context cache, and other open programs, then compare the total with the usable memory in the previous table. The site's calculator does this addition for you; just keep the logic in mind.
| Usage | What to load | Minimum requirements |
|---|---|---|
| Desktop assistant: summaries, email, translation | An 8B or 12B model in Q4, with a context of a few thousand tokens | 16 GB |
| Search your documents (RAG) | A 12B generation model, an embeddings model, and a reranker loaded together | 24 GB to leave some headroom |
| Code assistance for medium-sized projects | A 24B (Devstral Small 2) or a 20B to 26B MoE, long context | 24 GB, 32 GB if the context grows |
Regarding heat and noise, we found no public measurements of an M4 iMac under sustained LLM load, so we make no claims about its fan. The only reliable reference is actual usage: if generation runs for hours, monitor frequency and temperature in Activity Monitor and keep the context reasonable.
#Getting started
- 01Check the memoryCheck your iMac's memory in the Apple menu, then About This Mac. It determines the list of possible models.
- 02Install Ollama or LM StudioOllama accelerates computations on Apple GPUs via the Metal API. LM Studio provides a GUI well suited to large screens.
- 03Load a suitable modelChoose a model from the table based on your memory, in Q4_K_M, and launch it with ollama run.
- 04Monitor GPU usageUse ollama ps to verify that the model is loaded on the GPU, then use Activity Monitor to track memory pressure.
- 05Limit contextWith 16 GB, keep the context to a few thousand tokens to prevent the system from swapping.
#iMac M4 or Mac mini M4: the real criterion
The Mac mini M4 uses the same chip and the same bandwidth; with the same amount of memory, the iMac is therefore not slower. What changes is the integrated, non-replaceable display and the lack of a Pro variant: the iMac tops out at 32 GB and 120 GB/s, while the Mac mini M4 Pro reaches 273 GB/s according to the same reference discussion. If your goal is a 24B or larger model at a reasonable speed, the Mac mini M4 Pro's bandwidth doubles the throughput; if it is a desktop assistant with an 8 to 14B model, the iMac is equivalent.
| Criterion | iMac M4 | Mac mini M4 / M4 Pro |
|---|---|---|
| Bandwidth | 120 GB/s | 120 GB/s (M4); 273 GB/s (M4 Pro) |
| Maximum memory | 32 GB | View the Mac mini page |
| Screen | Built-in 24-inch 4.5K | Choose separately |
| Measured throughput 7B Q4_0 | 24,11 t/s | 24.11 t/s (M4); 49.64 t/s (M4 Pro) |
| Scalability | None, non-replaceable memory | Same for the memory, replaceable screen |
- Mac mini M4 and M4 Pro: which LLM for each memory size
- Mac mini M4 Pro hardware specifications
- iMac technical specifications, Apple
- llama.cpp performance on Apple Silicon, llama.cpp repository
- Ollama documentation: supported GPUs, including Metal
- ModelPiper: the GPU memory limit on Mac
#Frequently asked questions
Is the M4 iMac good for local AI?+
16, 24, or 32 GB on an iMac M4 for an LLM?+
What speed can you expect from an LLM on an M4 iMac?+
Why no iMac M4 Pro?+
Is the iMac M4 with 8 GPU cores sufficient for an LLM?+
Does Ollama use the iMac M4's GPU?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.