Beginner 3 minOllama

Install Ollama on macOS (Apple Silicon): guide 2026

Apple Apple Silicon Macs (M1, M2, M3, M4) are among the best consumer machines for local AI. Unified memory allows a 64 GB MacBook Pro to run 70B models that a €2,000 gaming PC can't even approach. Here's how to get started in 3 minutes.

By Mohamed Meguedmi·Update 2026-08-27·Tested on macOS 14+

#Why use a Mac for local AI?

On a typical PC, the GPU's VRAM is physically separate from the CPU's RAM. Do you have a RTX 4070 with 12 GB? You are limited to 12 GB for your models, regardless of your 64 GB of system RAM.

On a Apple Apple Silicon Mac, memory is unified: the GPU, CPU, and Neural Engine share the same RAM. A 64 GB MacBook Pro M3 Max can allocate 48 GB to a model—enough to load a large MoE such as Qwen 3.6 35B-A3B at full precision, or run several models at once, effortlessly.

→
Rule of thumb
macOS reserves about 25% of RAM for the system. On a 16 GB Mac, expect ~12 GB usable. On 32 GB: ~24 GB. On 64 GB: ~48 GB.

#Prerequisites

The Mac Kit

Ollama runs on your Mac. The Mac kit turns it into a reliable service — automatic startup, networking, logs (ch. 7) —, connects Open WebUI to it (ch. 8), and explains how unified memory affects model selection (ch. 2).

  • Lifetime online access
  • PDF + files
  • Lifetime updates
macOS 12 Monterey or later
Sonoma (14) or Sequoia (15) recommended for the latest Metal optimizations.
Mac Apple Silicon
M1, M2, M3, or M4—Pro, Max, Ultra, whatever. Intel chips are not officially supported (slow).
16 GB of RAM minimum
8 GB is enough for a 3B, but it's tight. 16 GB opens up 8–9B models, 32 GB supports 30B MoE models (such as Qwen 3.6 35B-A3B), and 64 GB supports large models at full precision.
~20 GB of disk space
For 3–4 models. Downloads go into ~/.ollama/models.

#1. Installation

Two solid methods: the official .dmg or Homebrew. Developers will prefer Homebrew; everyone else will choose the simpler option.

#Option A — .dmg application

Official download
https://ollama.com/download/mac
  1. 01
    Open the downloaded .dmg
    Drag Ollama.app to your Applications folder, like any other Mac app.
  2. 02
    Launch Ollama.app once
    On first launch, macOS asks for confirmation (“app downloaded from the internet”). Click Open.
  3. 03
    Authorize the command-line tool
    Ollama offers to install the ollama command. Accept—it will ask for an admin password.
  4. 04
    The llama icon appears in the menu bar
    That's the daemon. You can close the welcome window; it will keep running.

#Option B — Homebrew

Terminal
brew install --cask ollama

Faster to update. Advantage: brew upgrade ollama when a new version is released.

#Verification

Terminal
ollama --version

#2. Your first model

On a 16 GB Mac, start with Qwen 3.5 9B or Granite 4.2 8B. With 32 GB+, you can target Qwen 3.8 27B directly or a fast MoE such as Qwen 3.6 35B-A3B.

Qwen 3.5 9B (all configurations ≥16 GB)
ollama run qwen3.5:9b
Granite 4.2 8B (≥16 GB comfortable)
ollama run granite4.2:8b
Qwen 3.8 27B (≥32 GB recommended)
ollama run qwen3.8:27b
Qwen 3.6 35B-A3B MoE (≥32 GB, comfortable on 64 GB)
ollama run qwen3.6:35b
i
First download
The model arrives through the Ollama CDN (several GB). Allow 2 to 10 minutes depending on your connection. Subsequent launches are instantaneous.

#3. Understanding unified memory

macOS dynamically allocates RAM among apps. When Ollama loads a model, the system automatically moves other things to swap if needed. You can therefore load a larger model than usual, at the cost of an overall slowdown.

To monitor this in real time while a model is running:

Memory monitor
top -o mem
!
Excessive swapping = slowness
If you see more than 10 GB of active swap while Ollama is running, your model is too large for your Mac. Drop down one size or choose more aggressive quantization (Q3 instead of Q4).

#4. Verify Metal acceleration

Metal is Apple’s equivalent of CUDA: the API that lets Ollama run computations on the Mac’s integrated GPU. It’s enabled by default, but worth checking.

Loaded model status
ollama ps

In the PROCESSOR column, you should see 100% GPU. If you see CPU, something is wrong — either the model is unsupported or the Metal driver is too old.

Live GPU usage
sudo powermetrics --samplers gpu_power -i 500

#5. A ChatGPT-like interface

For daily use, the terminal gets tiring quickly. Three polished options on Mac:

Enchanted (free, App Store)
The cleanest design-wise. Native SwiftUI, connects to Ollama with no configuration.
Msty (free, .dmg)
More complete: multi-model support, tabs, saved prompts, built-in RAG.
Open WebUI via Docker
If you already have Docker, this is the most full-featured version (history, Markdown, extensions).
Enchanted in 1 line
open 'macappstore://apps.apple.com/app/enchanted-llm/id6474268307'

#6. Ollama at startup

By default, Ollama starts with your session if you installed it through the .dmg. With Homebrew, it is manual:

Auto-start via brew services
brew services start ollama

To disable it: brew services stop ollama. To check the status: brew services list.

#Troubleshooting

"ollama: command not found" after installation
The symbolic link was not created. Launch Ollama.app once and accept the offer to install the CLI.
The model crashes halfway through
Insufficient memory. Close Chrome/Electron and restart. Otherwise, use a smaller quantization.
Fans running full blast on MacBook
Normal during intensive inference. To limit it: use a smaller model, or plug into AC power (MacBooks throttle on battery).
Unable to delete Ollama
Stop the daemon: ollama stop. Then delete Ollama.app and the ~/.ollama directory (contains all downloaded models).
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.