Intermediate 11 minPerformance

Get the most out of a Apple Mac Silicon

Direct response

Optimizing an LLM on Apple Silicon means raising the memory limit macOS allows the GPU (sudo sysctl iogpu.wired_limit_mb), choosing a suitable model and quantization, quantizing the KV cache for long contexts, and selecting the right engine. No setting can exceed the chip's bandwidth: it caps generation speed.

A Mac isn't a PC with a GPU: its memory is shared, and the portion usable by the GPU can be configured. This page explains the GPU memory limit and how to change it, which models fit each memory size, KV cache, flash attention, and engine selection.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Optimizing an LLM on Apple Silicon: the four levers

To get the most out of a Mac Apple Silicon, four levers matter, in this order: the memory limit macOS allows the GPU to use, the model size and quantization, the KV cache and flash attention for long contexts, and the engine choice (Ollama, llama.cpp, MLX, or LM Studio). It all starts with one fact: unified memory is shared between the CPU and GPU, so total RAM serves as VRAM, but macOS caps the amount the GPU can address. Raising this limit enables larger models; the other levers prevent memory saturation or slow generation. None can make the machine faster than its bandwidth allows.

Unified memory
The CPU and GPU read the same data without copying. MLX summarizes it this way: arrays live in shared memory. So a 64 GB Mac can load a model that no 24 or 32 GB consumer graphics card can fit.
Efficiency
A Mac uses much less power than a PC equipped with a high-end graphics card, whose power draw exceeds 500 W at NVIDIA for RTX 5090: the Mac remains quiet and power-efficient for light workloads.
Ecosystem
The llama.cpp repository describes Apple silicon as a first-class citizen, optimized through ARM NEON, Accelerate, and Metal. Ollama, LM Studio, and MLX also use Metal.
Limitations
No CUDA, so libraries that require it won't run; a Mac's bandwidth (up to 546 GB/s on an M4 Max) remains below that of a high-end graphics card; fine-tuning is possible but slower.

#Mac VRAM: the GPU memory limit and iogpu.wired_limit_mb

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

On a Mac, there is no separate VRAM: “VRAM” is a portion of unified memory that macOS allows the GPU to wire. This portion is smaller than the total RAM. Public sources disagree on the default fraction, ranging from two-thirds to three-quarters depending on the case; don’t assume it—read it on your machine. MLX exposes the value selected by the system through device_info(), in the max_recommended_working_set_size field, and the MLX documentation says you can read this system limit in megabytes with sudo sysctl iogpu.wired_limit_mb.

The mlx-lm repository gives the safety rule: the value must exceed the model size in megabytes while remaining below the machine’s memory. Apple does not document this setting as a consumer-facing parameter; treat it as a risky system modification.

Rule of thumb for choosing the value (a cautious reserve for macOS, to be adjusted)
Mac memoryRecommended system reserveTarget GPU memoryValue of iogpu.wired_limit_mb
16 GB≈ 5 GB≈ 11 GB11264
24 GB≈ 6 GB≈ 18 GB18432
32 GB≈ 8 GB≈ 24 GB24576
64 GB≈ 8 to 10 GB≈ 54 to 56 GB55296 à 57344
128 GB≈ 16 GB≈ 112 GB114688

The formula is simple: value in MB = target memory in GB × 1,024. The reserve is a precautionary choice in this guide, not an Apple datum: it depends on the applications you leave open. If the machine slows down, writes to disk, or becomes unstable, reduce the value.

To read the limit macOS actually uses, one line of Python is enough if the mlx package is installed. The output contains the max_recommended_working_set_size field, expressed in bytes: divide by 1 073 741 824 to get gigabytes. Compare this figure with your model’s size before changing any setting: if the model already fits below it, raising the limit does nothing.

Terminal
python3 -c "import mlx.core as mx; print(mx.device_info())"
  1. 01
    Read the current value
    In Terminal, the sysctl iogpu.wired_limit_mb command displays the value. A zero generally indicates that macOS is applying its default limit.
  2. 02
    Calculate the New Value
    Take the model size plus its KV cache, add some headroom, then leave the system a reserve of at least 5 GB. Convert the result to megabytes.
  3. 03
    Apply the setting
    Run sudo sysctl iogpu.wired_limit_mb with the desired value. According to published guides, it takes effect without a reboot.
  4. 04
    Check
    Restart the model: the logs from Ollama or llama.cpp indicate the recommended working memory; with MLX, check device_info() again. Monitor memory pressure in Activity Monitor.
  5. 05
    Go back
    Set the value to 0 to return control to the system, or restart: the setting may not survive a restart; check it with sysctl.
Terminal
# Lire la valeur actuelle (0 = plafond par défaut du système)
sysctl iogpu.wired_limit_mb

# Exemple : 24 Go au GPU sur un Mac de 32 Go (24 x 1024)
sudo sysctl iogpu.wired_limit_mb=24576

# Retour au comportement par défaut
sudo sysctl iogpu.wired_limit_mb=0
!
What this setting does not do
It doesn't create memory; it allows the GPU to wire up more of it. A GPU that wires up too much leaves too little for the system, which starts swapping to disk: generation collapses. The mlx-lm repository also specifies that automatic wiring of models that occupy a lot of memory requires macOS 15 or later.

#Which models to use based on Mac memory

A model’s weight in Q4 follows the site’s guidelines: 3B is about 2 GB, 7–8B about 5 GB, 14B about 9 GB, 32B about 19 to 20 GB, and 70B about 40 GB. The KV cache and system headroom are added to that weight. The table cross-references these guidelines with Mac memory, before and after raising the GPU limit.

Realistic model class based on Mac memory (Q4_K_M)
MemoryWithout tuningWith the limit raised
16 GB7-9B, medium context12B, short context
24 GB12-14B24B in Q4 (≈ 14 GB), medium context
32 GB24B32B in Q4 (≈ 19-20 GB), short context
64 GB32B with long context70B in Q4 (≈ 40 GB), medium context
128 GB70B with long context70B with a very long context, or multiple loaded models

Maximum speed is calculated using chip bandwidth: bandwidth divided by weight in gigabytes. It is listed in each machine's specifications: see the MacBook Air, MacBook Pro, Mac mini, and Mac Studio guides, each of which includes its chip's figures.

A practical example: on a 32 GB Mac, a 32-billion-parameter model in Q4 weighs about 19 to 20 GB. With an 8,000-token context and an f16 KV cache, you need a few additional gigabytes, for about 23 to 24 GB total. If the default GPU limit is lower, part of the model moves to the CPU and speed drops: raising the limit to 24 or 26 GB solves the problem, provided you leave the system enough room to breathe. If memory pressure turns red, choose a 24-billion-parameter model instead.

#Quantization: why Q4 remains the right default on Mac

On a Mac, generation is limited by memory bandwidth: a smaller model is read faster. Q4_K_M therefore remains the default choice, offering the best balance of size, speed, and quality. Q5_K_M and Q6_K cost a few percentage points of speed for a modest quality gain; Q8_0 nearly doubles the weight of Q4 and therefore roughly halves the speed ceiling. 4-bit MLX formats play the same role in the MLX ecosystem. The quantization guide details the trade-off, so there is no need to repeat the numbers here.

#Metal and flash attention

Common engines use the Mac's GPU through Metal with no configuration. With Ollama, the ollama ps command indicates whether the model is loaded on the GPU. Flash attention reduces the memory used by long contexts: Ollama enables it automatically when the engine and hardware allow it, and the OLLAMA_FLASH_ATTENTION=1 variable forces it. In llama.cpp, the -fa option accepts on, off, or auto, with auto as the default; the -ngl option sets the number of layers placed on the GPU.

Terminal
# Ollama : forcer l'attention flash
export OLLAMA_FLASH_ATTENTION=1
ollama serve

# llama.cpp : toutes les couches sur le GPU, attention flash active
llama-server -m modele.gguf -ngl all -fa on -c 16384

#KV cache: the share of memory we forget

The KV cache stores vectors from every layer for each token in the context. Its size is twice the number of layers, times the number of KV heads, times their dimension, times the context length, times the number of bytes per value. Purely arithmetic example, with a typical architecture of 32 layers, 8 KV heads with dimension 128, and a context of 32 000 tokens in f16: 2 × 32 × 8 × 128 × 32 000 × 2 bytes, or approximately 4,2 GB, in addition to the weights. In q8_0, this cache drops to roughly half; in q4_0, to roughly a quarter, with a slight loss of precision.

Ollama quantizes the cache with the OLLAMA_KV_CACHE_TYPE variable (f16 by default); llama.cpp does it with the -ctk option for keys and -ctv for values. On a 32 GB Mac that allocates 24 GB to the GPU, saving a few gigabytes of cache can make the difference between a context that does not fit and a comfortable context. Ollama also chooses the default context based on memory: 4,000 tokens under 24 GiB.

Terminal
# Ollama
export OLLAMA_KV_CACHE_TYPE=q8_0
ollama serve

# llama.cpp
llama-server -m modele.gguf -ngl all -fa on -ctk q8_0 -ctv q8_0 -c 32768

#Ollama, LM Studio, llama.cpp, or MLX

Which engine for which need on Mac
EngineStrengthPreferred when
OllamaSimple local installation and APIDaily use, integration with other tools
LM StudioGraphical interface, MLX and GGUF modelsExplore models without a terminal
llama.cppFine-grained options, latest featuresYou want to tune GPU layers, KV cache, and server
MLX / mlx-lmApple framework, fine-tuning, shared memoryApple formats, lightweight training, Python scripts

The mlx-lm repository describes the package as a tool for text generation and model fine-tuning on Apple silicon, with support for quantized models. For comparing speeds between MLX and llama.cpp, refer to the dedicated guide; this page is limited to the shared settings.

#Battery and power-saving mode

On a MacBook, continuous generation keeps the GPU under constant load: battery life drops well below the web-browsing battery life advertised by Apple. This page does not give a precise duration because there is no source; it depends on the model and workload. macOS Low Power Mode reduces performance: reserve it for travel when battery life matters more than speed, and plug in the machine for any long session.

#Frequently asked questions

FAQ
How much VRAM does a Mac Apple Silicon have?+
No separate VRAM: the GPU uses unified memory. It can wire up a portion of the total RAM, less than 100%. Sources differ on the default fraction, between two-thirds and three-quarters, so read the system value with the MLX device_info command or Ollama logs, and check it as needed with iogpu.wired_limit_mb.
How can you increase a Mac's GPU memory with sysctl?+
Run sudo sysctl iogpu.wired_limit_mb followed by the value in megabytes, for example 24576 for 24 GB. The value must exceed the model size but remain below total memory, according to the mlx-lm repository. Leave at least 5 GB for the system. Set it back to 0 to return to the default behavior.
Does the iogpu.wired_limit_mb setting survive a reboot?+
Don't assume it: check the value with sysctl after every reboot. If it has returned to 0, you must reapply it. Apple does not document this setting; any method for applying it automatically at startup is at your own risk, and a value that is too high can make the machine unstable.
Should you use MLX or llama.cpp on a Mac?+
It depends on the use case. MLX is the Apple framework, suited to Apple formats and fine-tuning; llama.cpp offers fine-grained controls, while Ollama, which is built on it, offers simplicity. The site's comparison guide measures the speed gap; start with Ollama and switch only when a specific need arises.
Can a Mac replace a NVIDIA graphics card for LLMs?+
For model size, often yes: a 64 or 128 GB Mac can load models that a 24 or 32 GB graphics card cannot hold. For speed, a high-end card has greater bandwidth, and CUDA opens up more tools. The choice therefore depends on the target model.
Why is my model slow even though it fits in memory?+
Several possible causes: a model that is too large for the GPU limit is partially offloaded to the CPU (check with ollama ps), memory pressure triggers swapping, or the context is too long. Reduce the context, quantize the KV cache, and raise the GPU limit before changing models.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.