Install Ollama on macOS (Apple Silicon): guide 2026
Apple Apple Silicon Macs (M1, M2, M3, M4) are among the best consumer machines for local AI. Unified memory allows a 64 GB MacBook Pro to run 70B models that a €2,000 gaming PC can't even approach. Here's how to get started in 3 minutes.
#Why use a Mac for local AI?
On a typical PC, the GPU's VRAM is physically separate from the CPU's RAM. Do you have a RTX 4070 with 12 GB? You are limited to 12 GB for your models, regardless of your 64 GB of system RAM.
On a Apple Apple Silicon Mac, memory is unified: the GPU, CPU, and Neural Engine share the same RAM. A 64 GB MacBook Pro M3 Max can allocate 48 GB to a model—enough to load a large MoE such as Qwen 3.6 35B-A3B at full precision, or run several models at once, effortlessly.
#Prerequisites
Ollama runs on your Mac. The Mac kit turns it into a reliable service — automatic startup, networking, logs (ch. 7) —, connects Open WebUI to it (ch. 8), and explains how unified memory affects model selection (ch. 2).
- Lifetime online access
- PDF + files
- Lifetime updates
- macOS 12 Monterey or later
- Sonoma (14) or Sequoia (15) recommended for the latest Metal optimizations.
- Mac Apple Silicon
- M1, M2, M3, or M4—Pro, Max, Ultra, whatever. Intel chips are not officially supported (slow).
- 16 GB of RAM minimum
- 8 GB is enough for a 3B, but it's tight. 16 GB opens up 8–9B models, 32 GB supports 30B MoE models (such as Qwen 3.6 35B-A3B), and 64 GB supports large models at full precision.
- ~20 GB of disk space
- For 3–4 models. Downloads go into ~/.ollama/models.
#1. Installation
Two solid methods: the official .dmg or Homebrew. Developers will prefer Homebrew; everyone else will choose the simpler option.
#Option A — .dmg application
- 01Open the downloaded .dmgDrag Ollama.app to your Applications folder, like any other Mac app.
- 02Launch Ollama.app onceOn first launch, macOS asks for confirmation (“app downloaded from the internet”). Click Open.
- 03Authorize the command-line toolOllama offers to install the ollama command. Accept—it will ask for an admin password.
- 04The llama icon appears in the menu barThat's the daemon. You can close the welcome window; it will keep running.
#Option B — Homebrew
Faster to update. Advantage: brew upgrade ollama when a new version is released.
#Verification
#2. Your first model
On a 16 GB Mac, start with Qwen 3.5 9B or Granite 4.2 8B. With 32 GB+, you can target Qwen 3.8 27B directly or a fast MoE such as Qwen 3.6 35B-A3B.
#3. Understanding unified memory
macOS dynamically allocates RAM among apps. When Ollama loads a model, the system automatically moves other things to swap if needed. You can therefore load a larger model than usual, at the cost of an overall slowdown.
To monitor this in real time while a model is running:
#4. Verify Metal acceleration
Metal is Apple’s equivalent of CUDA: the API that lets Ollama run computations on the Mac’s integrated GPU. It’s enabled by default, but worth checking.
In the PROCESSOR column, you should see 100% GPU. If you see CPU, something is wrong — either the model is unsupported or the Metal driver is too old.
#5. A ChatGPT-like interface
For daily use, the terminal gets tiring quickly. Three polished options on Mac:
- Enchanted (free, App Store)
- The cleanest design-wise. Native SwiftUI, connects to Ollama with no configuration.
- Msty (free, .dmg)
- More complete: multi-model support, tabs, saved prompts, built-in RAG.
- Open WebUI via Docker
- If you already have Docker, this is the most full-featured version (history, Markdown, extensions).
#6. Ollama at startup
By default, Ollama starts with your session if you installed it through the .dmg. With Homebrew, it is manual:
To disable it: brew services stop ollama. To check the status: brew services list.
#Troubleshooting
- "ollama: command not found" after installation
- The symbolic link was not created. Launch Ollama.app once and accept the offer to install the CLI.
- The model crashes halfway through
- Insufficient memory. Close Chrome/Electron and restart. Otherwise, use a smaller quantization.
- Fans running full blast on MacBook
- Normal during intensive inference. To limit it: use a smaller model, or plug into AC power (MacBooks throttle on battery).
- Unable to delete Ollama
- Stop the daemon: ollama stop. Then delete Ollama.app and the ~/.ollama directory (contains all downloaded models).
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.