Beginner 11 miniMac

Which LLM on an iMac M4 (16 / 24 / 32 GB) ?

Direct response

The iMac M4 comfortably runs 8- to 14-billion-parameter models in Q4: its M4 chip provides 120 GB/s of memory bandwidth, and a public benchmark reports 24 tokens per second on a 7B in Q4_0 with the 10-core GPU. Unified memory (16, 24, or 32 GB) sets the maximum size: 16 GB for a 14B, 24 GB for a slow 24B, and 32 GB for an even slower 30B.

An iMac is primarily a display; for local AI, it is mainly an M4 chip and a quantity of unified memory that can no longer be changed. This guide quantifies what each configuration can run, at what speed according to public measurements, and when a Mac mini is the better purchase.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What an iMac M4 can do for local AI

The iMac M4 runs 8- to 14-billion-parameter models well and accepts larger models at reduced speed. Three figures summarize the essentials. The M4 chip delivers 120 GB/s of memory bandwidth according to the Apple specifications, implying a theoretical ceiling of about 31 tokens per second on a 7B in Q4_0. The public llama.cpp repository benchmark on an M4 with a 10-core GPU reports 24.11 tokens per second during generation, or 77% of that ceiling. Finally, unified memory ranges from 16 to 32 GB and is chosen at purchase: it alone determines the possible model size. What this iMac lacks is bandwidth, not capacity.

Chip
Apple M4; the Apple specification does not offer a Pro or Max variant for the iMac.
Bandwidth
120 GB/s, regardless of the configuration.
Unified memory
16 GB standard; 24 GB optional on all versions, 32 GB on the 4-port model.
Screen
24-inch 4.5K (4480 × 2520), which leaves little reason to use a larger GPU: memory is the limit.
Pricing
Not indicated here: Apple modifies them; see the site's hardware page.

#Real-world configurations: two models, not three

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The Apple page distinguishes two families. The 2-port model has an 8-core CPU and an 8-core GPU, with 16 GB of memory expandable to 24 GB. The 4-port model has a 10-core CPU and a 10-core GPU, with 16 GB expandable to 24 or 32 GB. The 120 GB/s bandwidth is identical. A common error, present in the older version of this page, is to assume that an 8 GB configuration still exists: the base memory is 16 GB. For an LLM, the key point is therefore to choose the 4-port model if you want 32 GB.

Useful iMac M4 configurations for an LLM (Apple sheet)
VersionGPUUnified memoryBandwidth
2 ports8 cores16 GB, 24 GB option120 GB/s
4 ports10 cores16 GB, with 24 and 32 GB options120 GB/s

For text generation, the difference between 8 and 10 GPU cores is probably small: speed depends on bandwidth, which is identical. No public measurements from the reference table exist for the 8-core variant, so this claim remains a hypothesis. If you're unsure, pay for memory rather than cores.

#16, 24, or 32 GB: the memory budget

On a Mac, memory is shared between the system and the GPU, and macOS gives the GPU only a fraction of it. The iogpu.wired_limit_mb setting defaults to 0, meaning the kernel derives the limit from the installed memory. On a 32 GB Mac measured by ModelPiper, Metal made 24.96 GiB available to the GPU, or 78% of memory. 75% is commonly cited. For an iMac, therefore, expect roughly 11 to 12 GB for a 16 GB model, 17 to 18 GB for 24 GB, and 24 to 25 GB for 32 GB: these are estimates, so verify them on your machine with the sysctl iogpu.wired_limit_mb command and the Metal tool described by ModelPiper.

Memory available to the model (estimated at 75%)
ConfigurationMemory available to the GPULargest reasonable model
16 GB≈ 11 to 12 GBA 14B in Q4 (≈ 9 GB), medium context
24 GB≈ 17 to 18 GBA 24B in Q4 (≈ 14 GB), short context
32 GB≈ 24 to 25 GBA 30B in Q4 (≈ 17 to 18 GB), medium context

To go beyond these limits, you can increase the share allocated to the GPU. The guide to optimizing Macs with Apple Silicon details this setting and its risks; we do not repeat it here.

#Which models for which memory capacity

The Q4 weights below come from the QuelLLM catalog. They assume the model fits in the usable memory shown in the previous table, including the context. The “smoothness” column applies the 120 GB/s ceiling to each model: these are calculated estimates, not measurements.

Models by memory configuration (Q4, weights only)
ModelWeightsMinimum requirementsExpected smoothness on M4
Llama 3.1 8B≈ 6 GB16 GBSmooth (about 15 tokens/s)
Gemma 4 12B≈ 7 GB16 GBFair (around 13 tokens/s)
Phi-4 14B≈ 9 GB16 GBDecent (about 10 tokens/s)
gpt-oss 20B≈ 13 GB24 GBVariable (MoE: faster than a dense 20B model)
Devstral Small 2 24B≈ 14 GB24 GBSlow (approximately 6 to 7 tokens/s)
Gemma 4 31B≈ 18 GB32 GBVery slow (about 5 tokens/s)

The key is not to confuse “fits” with “is comfortable.” A dense 24B model on an M4 base remains readable but slow, at a few words per second, whereas a mixture-of-experts (MoE) model such as gpt-oss 20B or Gemma 4 26B-A4B reads only a few billion parameters at each token, so it outpaces a comparably sized dense model. For a 24 or 32 GB iMac, an MoE is often the best choice.

A common trade-off on 24 GB is a 12B in Q8 (about 13 GB) or a 24B in Q4 (about 14 GB). Both use almost the same amount of memory, so they generate at roughly the same speed, around 7 tokens per second based on the bandwidth ceiling. The latter offers greater reasoning capacity, while the former stays closer to the original model. With equal bandwidth, model size in gigabytes determines throughput, not the number of parameters.

#Speed: bandwidth is the deciding factor

The only public reference measurement is the llama.cpp repository table on Apple Silicon, which tests a 7B LLaMA in Q4_0. For an M4 with a 10-core GPU and 120 GB/s, it reports 221,29 tokens per second for prompt processing and 24,11 for generation. An M4 Pro at 273 GB/s with a 16-core GPU reports 49,64 for generation, a little more than twice as much. This ratio is close to that of the memory bandwidths (2,3 times): a sign that generation depends on memory, not on the number of GPU cores.

Apple Silicon on LLaMA 7B Q4_0 (llama.cpp discussion)
ChipGPUBandwidthGeneration tg (t/s)
M4 (4-port iMac)10 cores120 GB/s24,11
M4 Pro (Mac mini M4 Pro)16 cores273 GB/s49,64

These measurements date back to an early version of llama.cpp; recent versions and MLX may do better. Treat them as an order of magnitude. The previous version of this page reported 28 to 31 tokens per second for a 9B in Q5: that figure would exceed the 120 GB/s ceiling for a model of this size, so we removed it.

#Three realistic use cases based on memory

A memory budget is easier to understand when applied to a use case. In each case, add the model weights, context cache, and other open programs, then compare the total with the usable memory in the previous table. The site's calculator does this addition for you; just keep the logic in mind.

Typical use cases and recommended configuration
UsageWhat to loadMinimum requirements
Desktop assistant: summaries, email, translationAn 8B or 12B model in Q4, with a context of a few thousand tokens16 GB
Search your documents (RAG)A 12B generation model, an embeddings model, and a reranker loaded together24 GB to leave some headroom
Code assistance for medium-sized projectsA 24B (Devstral Small 2) or a 20B to 26B MoE, long context24 GB, 32 GB if the context grows

Regarding heat and noise, we found no public measurements of an M4 iMac under sustained LLM load, so we make no claims about its fan. The only reliable reference is actual usage: if generation runs for hours, monitor frequency and temperature in Activity Monitor and keep the context reasonable.

#Getting started

  1. 01
    Check the memory
    Check your iMac's memory in the Apple menu, then About This Mac. It determines the list of possible models.
  2. 02
    Install Ollama or LM Studio
    Ollama accelerates computations on Apple GPUs via the Metal API. LM Studio provides a GUI well suited to large screens.
  3. 03
    Load a suitable model
    Choose a model from the table based on your memory, in Q4_K_M, and launch it with ollama run.
  4. 04
    Monitor GPU usage
    Use ollama ps to verify that the model is loaded on the GPU, then use Activity Monitor to track memory pressure.
  5. 05
    Limit context
    With 16 GB, keep the context to a few thousand tokens to prevent the system from swapping.
Terminal
ollama ps
sysctl iogpu.wired_limit_mb
sysctl hw.memsize

#iMac M4 or Mac mini M4: the real criterion

The Mac mini M4 uses the same chip and the same bandwidth; with the same amount of memory, the iMac is therefore not slower. What changes is the integrated, non-replaceable display and the lack of a Pro variant: the iMac tops out at 32 GB and 120 GB/s, while the Mac mini M4 Pro reaches 273 GB/s according to the same reference discussion. If your goal is a 24B or larger model at a reasonable speed, the Mac mini M4 Pro's bandwidth doubles the throughput; if it is a desktop assistant with an 8 to 14B model, the iMac is equivalent.

iMac M4 or Mac mini
CriterioniMac M4Mac mini M4 / M4 Pro
Bandwidth120 GB/s120 GB/s (M4); 273 GB/s (M4 Pro)
Maximum memory32 GBView the Mac mini page
ScreenBuilt-in 24-inch 4.5KChoose separately
Measured throughput 7B Q4_024,11 t/s24.11 t/s (M4); 49.64 t/s (M4 Pro)
ScalabilityNone, non-replaceable memorySame for the memory, replaceable screen
→
The deciding criterion
Choose the iMac if you want an all-in-one workstation and an 8B to 14B assistant. Choose a Mac mini M4 Pro if you're targeting a 24B or larger model, where bandwidth doubles throughput. Prices change: check them on the site's hardware pages.

#Frequently asked questions

FAQ
Is the M4 iMac good for local AI?+
Yes for 8- to 14-billion-parameter models: the M4 chip offers 120 GB/s of bandwidth, and a public measurement gives 24 tokens per second on a 7B in Q4_0. It slows noticeably on 24B models and larger because of bandwidth, not capacity.
16, 24, or 32 GB on an iMac M4 for an LLM?+
Sixteen GB is enough for an 8B to 14B model in Q4, with about 11 to 12 GB usable. Twenty-four GB supports 20B to 24B models, and thirty-two GB supports 30B models in Q4. Since memory cannot be replaced after purchase, aim for the largest model you plan to use, not the one you use today.
What speed can you expect from an LLM on an M4 iMac?+
On an M4 with a 10-core GPU, the llama.cpp table reports 24.11 tokens per second for generation with a 7B LLaMA in Q4_0. The bandwidth-to-weights calculation suggests about 15 tokens per second for a recent 8B model in Q4 and 6 to 7 for a dense 24B model: estimates to verify.
Why no iMac M4 Pro?+
The Apple page offers only the M4 chip for the iMac, in versions with 8- and 10-core GPUs. For higher bandwidth, move to a Mac mini M4 Pro (273 GB/s), whose generation is more than twice as fast in the reference table.
Is the iMac M4 with 8 GPU cores sufficient for an LLM?+
Probably yes: at 120 GB/s for all variants, generation depends little on the number of GPU cores. No llama.cpp benchmark data exists for the 8-core variant, so this is an assumption. Choose based on memory instead: 24 GB is the first truly useful tier.
Does Ollama use the iMac M4's GPU?+
Yes: Ollama's documentation states that it accelerates computations on Apple GPUs through the Metal API. Check with ollama ps that the model is loaded on the GPU, and monitor memory pressure in Activity Monitor if you're approaching your configuration's limit.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.