Get the most out of a Apple Mac Silicon
Optimizing an LLM on Apple Silicon means raising the memory limit macOS allows the GPU (sudo sysctl iogpu.wired_limit_mb), choosing a suitable model and quantization, quantizing the KV cache for long contexts, and selecting the right engine. No setting can exceed the chip's bandwidth: it caps generation speed.
A Mac isn't a PC with a GPU: its memory is shared, and the portion usable by the GPU can be configured. This page explains the GPU memory limit and how to change it, which models fit each memory size, KV cache, flash attention, and engine selection.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Optimizing an LLM on Apple Silicon: the four levers
To get the most out of a Mac Apple Silicon, four levers matter, in this order: the memory limit macOS allows the GPU to use, the model size and quantization, the KV cache and flash attention for long contexts, and the engine choice (Ollama, llama.cpp, MLX, or LM Studio). It all starts with one fact: unified memory is shared between the CPU and GPU, so total RAM serves as VRAM, but macOS caps the amount the GPU can address. Raising this limit enables larger models; the other levers prevent memory saturation or slow generation. None can make the machine faster than its bandwidth allows.
- Unified memory
- The CPU and GPU read the same data without copying. MLX summarizes it this way: arrays live in shared memory. So a 64 GB Mac can load a model that no 24 or 32 GB consumer graphics card can fit.
- Efficiency
- A Mac uses much less power than a PC equipped with a high-end graphics card, whose power draw exceeds 500 W at NVIDIA for RTX 5090: the Mac remains quiet and power-efficient for light workloads.
- Ecosystem
- The llama.cpp repository describes Apple silicon as a first-class citizen, optimized through ARM NEON, Accelerate, and Metal. Ollama, LM Studio, and MLX also use Metal.
- Limitations
- No CUDA, so libraries that require it won't run; a Mac's bandwidth (up to 546 GB/s on an M4 Max) remains below that of a high-end graphics card; fine-tuning is possible but slower.
#Mac VRAM: the GPU memory limit and iogpu.wired_limit_mb
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
On a Mac, there is no separate VRAM: “VRAM” is a portion of unified memory that macOS allows the GPU to wire. This portion is smaller than the total RAM. Public sources disagree on the default fraction, ranging from two-thirds to three-quarters depending on the case; don’t assume it—read it on your machine. MLX exposes the value selected by the system through device_info(), in the max_recommended_working_set_size field, and the MLX documentation says you can read this system limit in megabytes with sudo sysctl iogpu.wired_limit_mb.
The mlx-lm repository gives the safety rule: the value must exceed the model size in megabytes while remaining below the machine’s memory. Apple does not document this setting as a consumer-facing parameter; treat it as a risky system modification.
| Mac memory | Recommended system reserve | Target GPU memory | Value of iogpu.wired_limit_mb |
|---|---|---|---|
| 16 GB | ≈ 5 GB | ≈ 11 GB | 11264 |
| 24 GB | ≈ 6 GB | ≈ 18 GB | 18432 |
| 32 GB | ≈ 8 GB | ≈ 24 GB | 24576 |
| 64 GB | ≈ 8 to 10 GB | ≈ 54 to 56 GB | 55296 à 57344 |
| 128 GB | ≈ 16 GB | ≈ 112 GB | 114688 |
The formula is simple: value in MB = target memory in GB × 1,024. The reserve is a precautionary choice in this guide, not an Apple datum: it depends on the applications you leave open. If the machine slows down, writes to disk, or becomes unstable, reduce the value.
To read the limit macOS actually uses, one line of Python is enough if the mlx package is installed. The output contains the max_recommended_working_set_size field, expressed in bytes: divide by 1 073 741 824 to get gigabytes. Compare this figure with your model’s size before changing any setting: if the model already fits below it, raising the limit does nothing.
- 01Read the current valueIn Terminal, the sysctl iogpu.wired_limit_mb command displays the value. A zero generally indicates that macOS is applying its default limit.
- 02Calculate the New ValueTake the model size plus its KV cache, add some headroom, then leave the system a reserve of at least 5 GB. Convert the result to megabytes.
- 03Apply the settingRun sudo sysctl iogpu.wired_limit_mb with the desired value. According to published guides, it takes effect without a reboot.
- 04CheckRestart the model: the logs from Ollama or llama.cpp indicate the recommended working memory; with MLX, check device_info() again. Monitor memory pressure in Activity Monitor.
- 05Go backSet the value to 0 to return control to the system, or restart: the setting may not survive a restart; check it with sysctl.
#Which models to use based on Mac memory
A model’s weight in Q4 follows the site’s guidelines: 3B is about 2 GB, 7–8B about 5 GB, 14B about 9 GB, 32B about 19 to 20 GB, and 70B about 40 GB. The KV cache and system headroom are added to that weight. The table cross-references these guidelines with Mac memory, before and after raising the GPU limit.
| Memory | Without tuning | With the limit raised |
|---|---|---|
| 16 GB | 7-9B, medium context | 12B, short context |
| 24 GB | 12-14B | 24B in Q4 (≈ 14 GB), medium context |
| 32 GB | 24B | 32B in Q4 (≈ 19-20 GB), short context |
| 64 GB | 32B with long context | 70B in Q4 (≈ 40 GB), medium context |
| 128 GB | 70B with long context | 70B with a very long context, or multiple loaded models |
Maximum speed is calculated using chip bandwidth: bandwidth divided by weight in gigabytes. It is listed in each machine's specifications: see the MacBook Air, MacBook Pro, Mac mini, and Mac Studio guides, each of which includes its chip's figures.
A practical example: on a 32 GB Mac, a 32-billion-parameter model in Q4 weighs about 19 to 20 GB. With an 8,000-token context and an f16 KV cache, you need a few additional gigabytes, for about 23 to 24 GB total. If the default GPU limit is lower, part of the model moves to the CPU and speed drops: raising the limit to 24 or 26 GB solves the problem, provided you leave the system enough room to breathe. If memory pressure turns red, choose a 24-billion-parameter model instead.
#Quantization: why Q4 remains the right default on Mac
On a Mac, generation is limited by memory bandwidth: a smaller model is read faster. Q4_K_M therefore remains the default choice, offering the best balance of size, speed, and quality. Q5_K_M and Q6_K cost a few percentage points of speed for a modest quality gain; Q8_0 nearly doubles the weight of Q4 and therefore roughly halves the speed ceiling. 4-bit MLX formats play the same role in the MLX ecosystem. The quantization guide details the trade-off, so there is no need to repeat the numbers here.
#Metal and flash attention
Common engines use the Mac's GPU through Metal with no configuration. With Ollama, the ollama ps command indicates whether the model is loaded on the GPU. Flash attention reduces the memory used by long contexts: Ollama enables it automatically when the engine and hardware allow it, and the OLLAMA_FLASH_ATTENTION=1 variable forces it. In llama.cpp, the -fa option accepts on, off, or auto, with auto as the default; the -ngl option sets the number of layers placed on the GPU.
#KV cache: the share of memory we forget
The KV cache stores vectors from every layer for each token in the context. Its size is twice the number of layers, times the number of KV heads, times their dimension, times the context length, times the number of bytes per value. Purely arithmetic example, with a typical architecture of 32 layers, 8 KV heads with dimension 128, and a context of 32 000 tokens in f16: 2 × 32 × 8 × 128 × 32 000 × 2 bytes, or approximately 4,2 GB, in addition to the weights. In q8_0, this cache drops to roughly half; in q4_0, to roughly a quarter, with a slight loss of precision.
Ollama quantizes the cache with the OLLAMA_KV_CACHE_TYPE variable (f16 by default); llama.cpp does it with the -ctk option for keys and -ctv for values. On a 32 GB Mac that allocates 24 GB to the GPU, saving a few gigabytes of cache can make the difference between a context that does not fit and a comfortable context. Ollama also chooses the default context based on memory: 4,000 tokens under 24 GiB.
#Ollama, LM Studio, llama.cpp, or MLX
| Engine | Strength | Preferred when |
|---|---|---|
| Ollama | Simple local installation and API | Daily use, integration with other tools |
| LM Studio | Graphical interface, MLX and GGUF models | Explore models without a terminal |
| llama.cpp | Fine-grained options, latest features | You want to tune GPU layers, KV cache, and server |
| MLX / mlx-lm | Apple framework, fine-tuning, shared memory | Apple formats, lightweight training, Python scripts |
The mlx-lm repository describes the package as a tool for text generation and model fine-tuning on Apple silicon, with support for quantized models. For comparing speeds between MLX and llama.cpp, refer to the dedicated guide; this page is limited to the shared settings.
#Battery and power-saving mode
On a MacBook, continuous generation keeps the GPU under constant load: battery life drops well below the web-browsing battery life advertised by Apple. This page does not give a precise duration because there is no source; it depends on the model and workload. macOS Low Power Mode reduces performance: reserve it for travel when battery life matters more than speed, and plug in the machine for any long session.
- MLX vs. llama.cpp: speed comparison
- MacBook Pro M4 Pro and Max
- Mac Studio: up to 512 GB
- Compile llama.cpp with Metal
- Quantize the KV cache
- Site VRAM calculator
- Source: mlx-lm repository (wired memory limit)
- Source: MLX documentation, set_wired_limit
- Source: Ollama FAQ, flash attention, and KV cache
- Source: llama.cpp repository
#Frequently asked questions
How much VRAM does a Mac Apple Silicon have?+
How can you increase a Mac's GPU memory with sysctl?+
Does the iogpu.wired_limit_mb setting survive a reboot?+
Should you use MLX or llama.cpp on a Mac?+
Can a Mac replace a NVIDIA graphics card for LLMs?+
Why is my model slow even though it fits in memory?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.