The reference guide · 16 chapters · asset pack included · free updates for life
Local AI on your Mac,
100% Metal
The Mac Kit is THE reference guide for running local AI at full speed on Apple Silicon — unified memory, MLX vs GGUF, the right model for YOUR chip (M1 through M5/M6), and everything macOS gives you on top. For a Mac you just bought and for a 2020 Mac you never pushed to its limit.
Get the kit — $24$24 · one-time payment · free updates for life
What is The Mac Kit?
The Mac Kit is a 16-chapter reference guide and 21-file asset pack for running local AI at full speed on Apple Silicon. It's built for Mac users on any chip from M1 through M5/M6, from an 8GB MacBook Air to a 512GB Mac Studio, who already have or are setting up LM Studio or Ollama. You get the chip-by-RAM model table, MLX versus GGUF guidance, and launchd scripts for a Mac running as a server, with updates for life.
- 16 chapters and 92 pages, from unified memory basics to a 2026 Mac buying grid.
- 21-file asset pack: a chip-by-RAM model table, launchd scripts, and a Shortcuts export.
- Covers Apple Silicon chips M1 through M5/M6, with 8GB to 512GB of unified memory.
- 3 anchor machines: MacBook Air M2 16GB, Mac mini M4 24GB, Mac Studio M4 Max 64GB.
- Includes the sysctl commands to raise the GPU's unified-memory limit yourself.
- $24 one-time payment, updates for life included, 30-day money-back guarantee.
✓ This is for you if…
you have an Apple Silicon Mac (MacBook Air/Pro, Mac mini, Mac Studio, iMac — M1 through M5/M6) and you want to understand and use what's specific to it for local AI: its unified memory, MLX, the right model for your exact chip, real-world performance — whether you're starting out or already running Ollama or LM Studio without being sure you're getting the most out of it.
✗ This isn't for you if…
you're on a PC or Linux (that's The Local AI Kit, multi-OS) or you're looking for a copilot that codes inside your editor (that's The Local Copilot Kit — this guide gives Xcode a single pointer, never a dedicated chapter).
Why you can trust this
This guide is written and maintained by BestLLMfor — an independent, continuous benchmark pipeline tracking 239 locally-runnable models and 73 guides, run in the open on a published methodology. Every version, Ollama tag, license and command in this guide was checked against official sources (ollama.com, Hugging Face model cards, project changelogs and documentation) at the edition date, September 2026. Nothing is presented as measured on our own machines unless the text says so explicitly; performance figures point to the living pages of bestllmfor.com. What you read here is condensed and verified, not a random compilation of tutorials found online. Who's behind BestLLMfor →
The problem
An Apple Silicon Mac is not just a PC with a logo on it when it comes to local AI: unified memory shared with the GPU, MLX vs GGUF, a default memory ceiling you can raise once you know how — most guides treat your Mac like any other machine, and you leave half its capacity on the table.
The Mac-specific information is already out there, scattered across dozens of free guides — one per machine, one for MLX, one for unified memory. This guide is THE reference method: the chip × RAM model table assembled in one place, the launchd and Shortcuts scripts ready to paste, a method to measure your own real performance — not a pile of tutorials you have to cross-check yourself.
MLX native vs GGUF on your Mac
Same model, two formats — the choice most free guides explain worst. The honest comparison:
| MLX (Apple) | GGUF (llama.cpp/Ollama) | |
|---|---|---|
| Designed for unified memory from the ground up | Yes | Compatible, not native |
| Model availability | A good share, not everything | Nearly everything |
| Cross-platform (PC, Linux) | No — Apple Silicon only | Yes |
| Switching inside LM Studio | One GUI toggle | One GUI toggle |
Command line (mlx-lm) | Optional, to go further | Native Ollama CLI |
This guide tells you exactly when MLX is a real win, and when GGUF is still the right call (model missing in MLX, need for cross-platform compatibility) — never a dogma, a method.
What you get
- Your members area, for life — all 16 chapters, always the current edition, readable in your browser.
- The full guide as a PDF — the same content, yours forever, offline.
- The 21-file asset pack — the master GB-to-model table by chip (
.md+.csv), the unified-memory cheat-sheet (sysctlcommands + a LaunchDaemon), the MLX/GGUF decision tree, the launchd.plistfor Ollama as a service, an Open WebUI compose file tuned for Mac, a Shortcuts export, and more. - 3 anchoring profiles — MacBook Air M2 16GB (your current machine, no longer sold new — see Ch.15), Mac mini M4 24GB, Mac Studio M4 Max 64GB: you find your own setup in every chapter, never an example that doesn't apply to you.
- The "when it breaks on Mac" chapter — a decision tree by symptom (Metal not detected, your system unusable during generation, a model that refuses to load, thermal throttling).
- Living edition — every update (new Apple chips, new models, MLX moving forward) shows up in your members area, for life, at no extra cost.
Tools: LM Studio (MLX), Ollama, Open WebUI, mlx-lm. Apple Silicon exclusively, $0/month once it's set up.
The 16 chapters of the kit
92 pages, from your first 100% Metal LLM in under an hour to the 2026 buying grid. Nothing to skip: every chapter links to the next.
- Your first 100% Metal LLM, in one hour
- Unified memory: what it actually changes (and how to raise the GPU limit)
- MLX or GGUF: which to choose, and when it matters
- The reference table: which model for WHICH chip, what RAM
- Quantization and long context: calculating your real memory budget
- Real-world performance: tokens per second, prefill, heat, battery
- Ollama as a service: auto-start, network access, logs (launchd)
- Open WebUI on Mac: Docker Desktop, OrbStack, or native
- Local vision, natively accelerated by Metal
- English and multilingual models that actually run well on Apple Silicon
- Shortcuts, Automator, Spotlight: local AI inside macOS
- Turn your Mac mini into a household AI server
- Privacy: iCloud, Spotlight, telemetry — what actually stays local
- When it breaks on Mac: Metal, memory, swap
- Which Mac to buy for local AI in 2026
- This edition has a shelf life — now what?
Living edition: every update enriches this table of contents (a dated update log, Ch.16) — you never repurchase to get the next version.
Your first 100% Metal LLM in one hour, step by step
Whether you're starting from zero or already have a tool installed, the guide walks you to a real result:
1 · LM Studio or Ollama
Two commands if you're starting from zero (a full pointer to The Local AI Kit for setup), otherwise go straight to step 2.
2 · Confirm Metal is running
Activity Monitor's GPU tab should show your GPU at work — not a silent CPU fallback.
3 · The right model for your chip
Chapter 4's table tells you which one, based on your chip and your unified memory — no guessing.
4 · Go further
Unified memory, MLX, launchd, Shortcuts — it's all in the chapters that follow.
- Your members area, for life: all 16 chapters always current (full table of contents above)
- The PDF guide + the 21-file asset pack (the master table, launchd/Shortcuts scripts, checklists) to download
- The 3 anchoring profiles (Air, mini, Studio) and the Metal/memory/swap troubleshooting tree
- Every future edition, for life — they show up inside your members area automatically
- Instant access after payment, access key emailed to you
Stripe receipt emailed automatically; invoice on request by replying to the delivery email.
The honest truth
An Apple Silicon Mac has real strengths for local AI — and real limits no chip table can fix. A reference guide owes you that honest map:
By the end, you'll know how to
- Explain unified memory and raise, if needed, the memory allocated to the GPU (
iogpu.wired_limit_mb). - Choose between MLX and GGUF depending on the model, without unnecessary jargon, and run both when it helps.
- Find the right model for YOUR chip and YOUR unified memory in one read of the master table.
- Run Ollama as a real macOS service (launchd) and turn a Mac mini into a household server.
- Verify for yourself that nothing leaves your Mac (iCloud Drive, Spotlight) — not just take our word for it.
- Diagnose and fix the Mac-specific failures yourself (Metal, memory, swap).
Living edition
"Free updates for life" isn't a slogan — Apple hardware changes twice a year and MLX moves fast; this guide tracks the ecosystem for you, at no extra cost.
v2026.09 — the launch edition: 16 chapters, 3 anchoring profiles (MacBook Air M2 16GB, Mac mini M4 24GB, Mac Studio M4 Max 64GB), a 21-file asset pack, the master GB-to-model table by chip.
Future editions (new Apple chips, new models, MLX moving forward) get added here, dated, and show up in your members area at no extra cost.
Frequently asked questions
What's the difference with The Local AI Kit?
The Local AI Kit sets up a private AI assistant on a Mac OR a PC, for a general, multi-OS audience — it stays deliberately broad. This guide assumes a Mac is already in the picture and goes deep on everything specific to Apple Silicon: unified memory and its GPU limit, MLX vs GGUF, the model table by chip AND by RAM tier (M1 through M5/M6, 8 to 512GB), a method to measure your own real performance, and what macOS gives you on top (Shortcuts, launchd, Spotlight, Apple Intelligence). The two complement each other; they don't overlap.
Does this only work on Apple Silicon Macs?
Yes, exclusively — a Mac with an M1 through M5/M6 chip (MacBook Air, MacBook Pro, Mac mini, Mac Studio, iMac). An Intel Mac has no unified memory and no MLX, so most of this guide won't apply. If you're on a PC, or you're not sure which Mac to buy yet, start with The Local AI Kit or Chapter 15 of this guide (a buying decision grid, no affiliate links).
Do I need to be a developer to follow this guide?
No. This guide assumes you know how to install an app and drag a file — the Terminal commands (unified memory in Ch.2, launchd in Ch.7) are explained the first time they appear, never assumed. A few sections go further than a general audience needs (mlx-lm on the command line) — always framed as "to go further," never a requirement to keep reading.
How long until I have a result?
Your first LLM with Metal acceleration confirmed active, in under an hour — whether you're starting from zero (two commands, with a pointer to The Local AI Kit for the full install) or you already have LM Studio or Ollama and just want to confirm Metal is really running, not a silent CPU fallback.
What's the difference with the free Mac guides on the site?
The free guides on bestllmfor.com cover each machine or each topic separately (one guide per Mac, MLX vs llama.cpp on one side, memory tuning on the other). This guide assembles the complete path, ships the chip × RAM model table that no single free guide keeps current and complete, gives you the ready-to-paste launchd and Shortcuts scripts, and stays current for life against Apple hardware that changes twice a year — it's the assembly and the assets you pay for, not the raw information.
Is it really free once it's set up?
Yes: the model runs on your machine, $0/month, no account, no limit. The guide itself (the book + the asset pack + a lifetime members area) is a one-time $24 purchase.
Can I really turn my Mac mini into a home AI server?
Yes — Chapter 12 covers exactly that: keeping it from sleeping, running Ollama as a real service (launchd, Ch.7) that survives a restart, and reaching it from the other devices in your home, with an honest note on remote-access options (a personal VPN like Tailscale) — never dressed up as an enterprise security checklist.
Do my conversations and models stay private on my Mac?
That's the whole point of local — nothing leaves once the model is loaded. Chapter 13 goes further than a generic checklist: it specifically checks whether iCloud Drive is silently syncing your models or conversations folder, and what Spotlight indexes around them.
What if it doesn't work out for me after I buy?
30-day money-back guarantee — one email, no questions asked.
How is it delivered, and for how long?
Instant access after payment: your lifetime members area (always the current edition) plus the PDF and the asset-pack zip to download, yours forever. Future updates (new Apple chips, new models) show up in your members area at no extra cost.
Can I get a receipt or an invoice?
Yes, automatically. As soon as you pay (via Stripe) you get an emailed Stripe receipt. A formal invoice is available on request — just reply to your delivery email.
What if I want more than one BestLLMfor kit?
All the kits, for life bundles every current and future kit for $49. If you've already bought one kit at $24, upgrading to the bundle costs $25, not $49 — the discount applies automatically.
Can a Mac mini run a local LLM well?
Yes -- a Mac mini's unified memory is shared with its GPU, so even the base 16GB configuration comfortably runs a mid-size quantized model, and the higher-RAM M4 Pro configurations handle much larger ones. The main limits are no memory upgrade after purchase and needing to keep it awake as a server, both covered with real settings in this guide.
What is the best local LLM to run on a Mac?
There's no single best model for Mac -- it depends on your chip and how much unified memory you have, from an 8GB MacBook Air up to a 512GB Mac Studio. This guide's chip-by-RAM table maps a specific recommended model to your exact configuration instead of one generic answer that ignores your hardware.
Do I need MLX or does GGUF work fine on Apple Silicon?
GGUF works fine on Apple Silicon through Ollama or llama.cpp and covers nearly every model released; MLX is Apple's own format, built for unified memory from the ground up, but a smaller share of models ship in it. This guide gives you a decision tree instead of a fixed rule: which one wins depends on whether your model exists in MLX and whether you need cross-platform compatibility.
Run your Mac at full speed.
THE reference guide, from your first Metal-accelerated LLM to a household AI server on your Mac mini — unified memory, MLX, the right model for your chip. Today, not "someday."
Get the kit — $24