Intermediate 12 minQwen

Qwen3 32B on Mac M1/M2: memory and settings MLX

Direct response

Qwen3 32B fits in 18.4 GB of RAM in the official 4-bit MLX version, plus the context used: plan on at least 24 to 32 GB of unified memory for comfortable use without swap. A base M1 or M2 Mac (8 or 16 GB) isn't enough; you need a Pro, Max, or Ultra chip.

Qwen3 32B is a dense model with 32.8 billion parameters, with a native context of 32,768 tokens extensible to 131,072 via YaRN. On a Mac, the question is not just “will it load?” but “with how much memory actually free, what quantization, and what context?” This guide details the verifiable figures for an M1 or M2 Mac, without assuming compatibility that does not exist on the more modest configurations.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#Unified memory actually required

The starting point is the weight of the quantized model, not the parameter count. The 4-bit MLX version published by the mlx-community weighs 18.4 GB on disk and in memory when loaded. On a Mac, unified memory is shared between the system, open applications, and the model: reserving only 18.4 GB of total RAM leaves no headroom for the system or the conversation context.

The official Qwen fact sheet specifies the architecture behind this model: 64 layers and grouped-query attention (GQA) with 64 query heads but only 8 key and value heads. In practice, fewer KV heads mean a more compact KV cache per context token than standard multi-head attention with equal numbers of heads—one reason a dense model with 32.8 billion parameters remains practical in unified memory on a Mac, provided the model's weights themselves already fit in the available RAM.

Unified memory required for Qwen3 32B
ConfigurationFeasibilityNote
16 GB of unified memoryNot realisticThe model alone (18.4 GB in 4-bit) already exceeds the total RAM.
24 GB of unified memoryPossible, short contextLittle headroom left for the system and KV cache; reserve it for occasional use.
32 GB of unified memoryComfortableEnough headroom for a medium-sized context without frequent swapping.
64 GB or moreComfortableAllows a long context while keeping other applications open.

#Which M1/M2 chip for 32B?

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A “basic” M1 or M2 MacBook Air or Pro comes with 8 or 16 GB of unified memory: insufficient for Qwen3 32B, even in 4-bit. You need to target an M1 Pro/Max/Ultra or M2 Pro/Max/Ultra chip configured with at least 32 GB of unified memory to have reasonable headroom once the system and context are taken into account. This isn't about GPU compute power; it's about the amount of physical memory in the machine, which can't be increased after purchase.

!
Objection: “my machine shows 16 GB, so it should work in swap”
Technically, macOS can use the SSD as auxiliary memory (swap), but SSD throughput is nowhere near that of unified memory: a model that swaps heavily becomes unusable in practice, with response times measured in minutes rather than seconds. This is not a viable option for regular use.

#MLX or llama.cpp/Ollama?

MLX is the framework developed to leverage the unified memory architecture of Apple Silicon; llama.cpp (and Ollama, which builds on it) remains a solid alternative, with a broader GGUF quantization ecosystem and a simpler learning curve for anyone already familiar with Ollama on other machines. For Qwen3 32B in particular, the official 4-bit MLX version from the mlx-community is directly comparable in size (18.4 GB) to equivalent GGUF quantizations: the choice depends mainly on which tool you already know, not on a difference in feasibility.

Choose MLX
If you're already in the Python/Hugging Face ecosystem and want the closest conversion to the Apple Silicon architecture.
Choose Ollama/llama.cpp
If you want a single command, a standard local server, and compatibility with the same tools as on PC/Linux.

#Install with MLX, step by step

  1. 01
    Check available memory
    In Activity Monitor (the Memory tab), note the memory that is actually free before launching anything, not just the total RAM advertised by the machine.
  2. 02
    Install mlx-lm
    Install the Python mlx-lm package in a dedicated environment rather than globally to avoid version conflicts.
  3. 03
    Download the 4-bit version
    Get mlx-community/Qwen3-32B-4bit instead of an unquantized version to limit the memory footprint during loading.
  4. 04
    Start with a reduced context
    Start with a short context (4,000 to 8,000 tokens) to verify that the model loads and responds, then gradually increase it based on the remaining memory you observe.

#The macOS GPU memory setting (unofficial)

On Apple Silicon, macOS reserves part of unified memory from the GPU by default: Metal exposes a recommended working-set size of about 75% of total RAM. A system setting, iogpu.wired_limit_mb, technically allows this limit to be raised. However, this is not an officially documented setting by Apple: it is a practice known to the technical community, unsupported, and liable to change or break in a future macOS update.

Raise the GPU memory ceiling (unofficial, at your own risk)
sudo sysctl iogpu.wired_limit_mb=26624

The value is in megabytes and must remain below physical memory: allow for total RAM minus 6 to 8 GB for macOS. The example above targets a 32 GB Mac (26,624 MB, or 26 GB for the GPU); on a 64 GB Mac, 57,344 MB leaves the same margin. Notes circulating with 122,880 MB target a 128 GB machine: copied as-is onto a 32 GB M1 or M2, this value would exceed the installed memory.

i
Keep in mind
This setting returns to its default value after a restart unless manually persisted, and Apple provides no stability guarantee for this configuration. Only use it if you're prepared to monitor system memory pressure and roll back if the system becomes unstable.

#Reasoning mode: an indirect memory cost

Qwen3-32B offers a switch between a thinking mode, which generates internal reasoning before the final answer, and a direct mode. This choice is made through the enable_thinking parameter or the /think and /no_think tags, depending on the tool used. On a machine at the limit of its memory, thinking mode has an indirect but real effect: every reasoning token generated before the final answer is added to the current conversation context, and therefore to the KV cache, which grows with it. On a Mac with 24-32 GB, disabling thinking mode for simple questions limits context growth and reduces the risk of memory saturation during a long session.

→
Objection: “reasoning mode improves quality, so why disable it?”
This is not about disabling it systematically, but about enabling it when the task warrants it (multistep reasoning, computation, complex code) and staying in direct mode for short exchanges, where the quality gain is marginal compared with the token and memory cost.

#Swap errors and signs of saturation

The clearest symptom of insufficient memory is not an explicit error message, but a sudden slowdown: the entire system becomes slow, not just the model, because macOS starts writing swap data to the SSD. In Activity Monitor, the memory pressure indicator turning orange or red while the model loads is the signal to watch before even waiting for the first response.

Orange/red memory pressure during loading
Reduce the requested context size, or switch to a lighter quantization if available.
The model loads, but the first response takes several minutes
Sign of active swapping: close memory-intensive applications before restarting.
Immediate loading failure
The machine’s total unified memory is probably insufficient for this model size, including quantization.

On the command line, the vm_stat command (number of swapped pages, “Pageouts”) or memory_pressure -Q provide a more precise diagnosis than Activity Monitor’s visual indicator alone. This helps confirm that a slowdown is actually caused by memory saturation rather than another bottleneck (a full disk or thermal throttling during prolonged use). An increasing number of pageouts while the model is generating is the most reliable sign that unified memory is insufficient for the current configuration, including the context.

#Long context: the real memory cost

Qwen3-32B's native 32,768-token context, extendable to 131,072 via YaRN, is not free in memory: each context token actually used feeds a cache (KV cache) that grows with the conversation. On a machine at the limit (24 to 32 GB), favoring a moderate context (8,000 to 16,000 tokens) leaves a safety margin; reserve the extension to 131,072 tokens for machines with 64 GB or more of unified memory, where the margin after loading the model remains comfortable.

#Recommended settings by mode

Choosing sampling parameters is not neutral on a memory-limited machine: a poorly adjusted setting can unnecessarily lengthen responses, and therefore the context and KV cache. Qwen publishes separate settings for each active mode, which you should pass explicitly in your MLX client or in Ollama's generation settings instead of keeping another model's generic defaults.

Recommended generation settings from Qwen by mode
ParameterReasoning mode (enable_thinking=True)Direct mode (enable_thinking=False)
Temperature0,60,7
Top P0,950,8
Top K2020
Min P00

Qwen also recommends planning for an output length of up to 32 768 tokens for most complex queries, and adjusting presence_penalty between 0 and 2 to limit repetitions if they appear—while keeping in mind that too high a value can cause language mixing in the response. On a Mac with tight memory, such a high maximum output length should remain theoretical: in practice, stopping generation as soon as the answer addresses the question prevents the KV cache from growing for no reason.

Frequently asked questions
Can an 8 GB Mac M1 run Qwen3 32B?+
No. The 4-bit version of the model already weighs 18.4 GB, more than the total memory of an 8 or 16 GB Mac, even before accounting for the system and context. You need at least a 24-32 GB unified-memory configuration, ideally a Pro, Max, or Ultra chip with 32 GB or more for comfortable use.
Should you choose MLX or Ollama for Qwen3 32B on Mac?+
Both work with comparable file sizes (around 18–20 GB in 4-bit). MLX is designed natively for the unified memory architecture of Apple Silicon; Ollama/llama.cpp offers a simpler experience if you already use it on other machines. The choice does not change the minimum memory required, only the tooling around the model.
Is the iogpu.wired_limit_mb setting safe?+
This is not an officially documented setting by Apple: it works in practice and is described by the technical community, but remains unsupported, resets to its default value after a restart, and may be changed by a future macOS update. Use it while monitoring the memory pressure indicator in Activity Monitor, ready to revert the change.
Which context should you use on a 32 GB machine?+
A moderate context, between 8,000 and 16,000 tokens, leaves enough safety margin after loading the model in 4-bit (18.4 GB) plus the system. Reserve the extended context of 131,072 tokens via YaRN for machines with 64 GB or more, where the KV cache has enough room to grow without saturating memory.
What generation settings should you use for Qwen3-32B?+
Qwen recommends temperature=0.6, top_p=0.95, top_k=20 in reasoning mode, and temperature=0.7, top_p=0.8, top_k=20 in direct mode, with min_p=0 in both cases. These values balance quality and response length; on a Mac with tight memory, they also prevent unnecessarily long generations that would increase the KV cache.
Why does reasoning mode use more memory?+
Reasoning mode generates internal reasoning before the final answer, and each token in that reasoning is added to the context of the current conversation—and therefore to the KV cache, which grows with it. On a Mac with 24–32 GB, disabling it for simple questions limits this growth and reduces the risk of running out of memory during a long session.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.