Qwen3 32B on Mac M1/M2: memory and settings MLX
Qwen3 32B fits in 18.4 GB of RAM in the official 4-bit MLX version, plus the context used: plan on at least 24 to 32 GB of unified memory for comfortable use without swap. A base M1 or M2 Mac (8 or 16 GB) isn't enough; you need a Pro, Max, or Ultra chip.
Qwen3 32B is a dense model with 32.8 billion parameters, with a native context of 32,768 tokens extensible to 131,072 via YaRN. On a Mac, the question is not just “will it load?” but “with how much memory actually free, what quantization, and what context?” This guide details the verifiable figures for an M1 or M2 Mac, without assuming compatibility that does not exist on the more modest configurations.
#Unified memory actually required
The starting point is the weight of the quantized model, not the parameter count. The 4-bit MLX version published by the mlx-community weighs 18.4 GB on disk and in memory when loaded. On a Mac, unified memory is shared between the system, open applications, and the model: reserving only 18.4 GB of total RAM leaves no headroom for the system or the conversation context.
The official Qwen fact sheet specifies the architecture behind this model: 64 layers and grouped-query attention (GQA) with 64 query heads but only 8 key and value heads. In practice, fewer KV heads mean a more compact KV cache per context token than standard multi-head attention with equal numbers of heads—one reason a dense model with 32.8 billion parameters remains practical in unified memory on a Mac, provided the model's weights themselves already fit in the available RAM.
| Configuration | Feasibility | Note |
|---|---|---|
| 16 GB of unified memory | Not realistic | The model alone (18.4 GB in 4-bit) already exceeds the total RAM. |
| 24 GB of unified memory | Possible, short context | Little headroom left for the system and KV cache; reserve it for occasional use. |
| 32 GB of unified memory | Comfortable | Enough headroom for a medium-sized context without frequent swapping. |
| 64 GB or more | Comfortable | Allows a long context while keeping other applications open. |
#Which M1/M2 chip for 32B?
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
A “basic” M1 or M2 MacBook Air or Pro comes with 8 or 16 GB of unified memory: insufficient for Qwen3 32B, even in 4-bit. You need to target an M1 Pro/Max/Ultra or M2 Pro/Max/Ultra chip configured with at least 32 GB of unified memory to have reasonable headroom once the system and context are taken into account. This isn't about GPU compute power; it's about the amount of physical memory in the machine, which can't be increased after purchase.
#MLX or llama.cpp/Ollama?
MLX is the framework developed to leverage the unified memory architecture of Apple Silicon; llama.cpp (and Ollama, which builds on it) remains a solid alternative, with a broader GGUF quantization ecosystem and a simpler learning curve for anyone already familiar with Ollama on other machines. For Qwen3 32B in particular, the official 4-bit MLX version from the mlx-community is directly comparable in size (18.4 GB) to equivalent GGUF quantizations: the choice depends mainly on which tool you already know, not on a difference in feasibility.
- Choose MLX
- If you're already in the Python/Hugging Face ecosystem and want the closest conversion to the Apple Silicon architecture.
- Choose Ollama/llama.cpp
- If you want a single command, a standard local server, and compatibility with the same tools as on PC/Linux.
#Install with MLX, step by step
- 01Check available memoryIn Activity Monitor (the Memory tab), note the memory that is actually free before launching anything, not just the total RAM advertised by the machine.
- 02Install mlx-lmInstall the Python mlx-lm package in a dedicated environment rather than globally to avoid version conflicts.
- 03Download the 4-bit versionGet mlx-community/Qwen3-32B-4bit instead of an unquantized version to limit the memory footprint during loading.
- 04Start with a reduced contextStart with a short context (4,000 to 8,000 tokens) to verify that the model loads and responds, then gradually increase it based on the remaining memory you observe.
#The macOS GPU memory setting (unofficial)
On Apple Silicon, macOS reserves part of unified memory from the GPU by default: Metal exposes a recommended working-set size of about 75% of total RAM. A system setting, iogpu.wired_limit_mb, technically allows this limit to be raised. However, this is not an officially documented setting by Apple: it is a practice known to the technical community, unsupported, and liable to change or break in a future macOS update.
The value is in megabytes and must remain below physical memory: allow for total RAM minus 6 to 8 GB for macOS. The example above targets a 32 GB Mac (26,624 MB, or 26 GB for the GPU); on a 64 GB Mac, 57,344 MB leaves the same margin. Notes circulating with 122,880 MB target a 128 GB machine: copied as-is onto a 32 GB M1 or M2, this value would exceed the installed memory.
#Reasoning mode: an indirect memory cost
Qwen3-32B offers a switch between a thinking mode, which generates internal reasoning before the final answer, and a direct mode. This choice is made through the enable_thinking parameter or the /think and /no_think tags, depending on the tool used. On a machine at the limit of its memory, thinking mode has an indirect but real effect: every reasoning token generated before the final answer is added to the current conversation context, and therefore to the KV cache, which grows with it. On a Mac with 24-32 GB, disabling thinking mode for simple questions limits context growth and reduces the risk of memory saturation during a long session.
#Swap errors and signs of saturation
The clearest symptom of insufficient memory is not an explicit error message, but a sudden slowdown: the entire system becomes slow, not just the model, because macOS starts writing swap data to the SSD. In Activity Monitor, the memory pressure indicator turning orange or red while the model loads is the signal to watch before even waiting for the first response.
- Orange/red memory pressure during loading
- Reduce the requested context size, or switch to a lighter quantization if available.
- The model loads, but the first response takes several minutes
- Sign of active swapping: close memory-intensive applications before restarting.
- Immediate loading failure
- The machine’s total unified memory is probably insufficient for this model size, including quantization.
On the command line, the vm_stat command (number of swapped pages, “Pageouts”) or memory_pressure -Q provide a more precise diagnosis than Activity Monitor’s visual indicator alone. This helps confirm that a slowdown is actually caused by memory saturation rather than another bottleneck (a full disk or thermal throttling during prolonged use). An increasing number of pageouts while the model is generating is the most reliable sign that unified memory is insufficient for the current configuration, including the context.
#Long context: the real memory cost
Qwen3-32B's native 32,768-token context, extendable to 131,072 via YaRN, is not free in memory: each context token actually used feeds a cache (KV cache) that grows with the conversation. On a machine at the limit (24 to 32 GB), favoring a moderate context (8,000 to 16,000 tokens) leaves a safety margin; reserve the extension to 131,072 tokens for machines with 64 GB or more of unified memory, where the margin after loading the model remains comfortable.
#Recommended settings by mode
Choosing sampling parameters is not neutral on a memory-limited machine: a poorly adjusted setting can unnecessarily lengthen responses, and therefore the context and KV cache. Qwen publishes separate settings for each active mode, which you should pass explicitly in your MLX client or in Ollama's generation settings instead of keeping another model's generic defaults.
| Parameter | Reasoning mode (enable_thinking=True) | Direct mode (enable_thinking=False) |
|---|---|---|
| Temperature | 0,6 | 0,7 |
| Top P | 0,95 | 0,8 |
| Top K | 20 | 20 |
| Min P | 0 | 0 |
Qwen also recommends planning for an output length of up to 32 768 tokens for most complex queries, and adjusting presence_penalty between 0 and 2 to limit repetitions if they appear—while keeping in mind that too high a value can cause language mixing in the response. On a Mac with tight memory, such a high maximum output length should remain theoretical: in practice, stopping generation as soon as the answer addresses the question prevents the KV cache from growing for no reason.
- MLX vs. llama.cpp on Mac: benchmark comparison
- Choose a model based on available RAM
- Install an LLM locally, step by step
- Quantize the KV cache to save memory with long contexts
- Source: official Qwen3-32B fact sheet (Hugging Face)
- Source: official 4-bit MLX version from mlx-community
- Source: iogpu.wired_limit_mb setting (unofficial Apple)
Can an 8 GB Mac M1 run Qwen3 32B?+
Should you choose MLX or Ollama for Qwen3 32B on Mac?+
Is the iogpu.wired_limit_mb setting safe?+
Which context should you use on a 32 GB machine?+
What generation settings should you use for Qwen3-32B?+
Why does reasoning mode use more memory?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.