BestLLMfor Your hardware. Your LLM. Your call.
The Local Copilot Kit APIOpen data Find my LLM
All models · 100% local, 100% private

Which LLM runs on
your machine?

Tell us what's under the hood. We'll tell you what runs, how fast, and how to install it — step by step, in plain English.

~/configurator —— loading…
Your data stays in your browser. Nothing is sent.
Free
No signup
Open data
CC BY 4.0
Independent
No tracker
Real benchmarks
Daily updates
Build your own local coding copilot — the reference kit. From first install to an agent that codes inside your editor. Get the kit →
What you get

Built around your decision, not vendor benchmarks.

Four practical tools that answer the questions you actually have when picking an LLM.

01 RANKINGS

Hardware-matched rankings

Best local LLM for RTX 4090, RTX 5090, Mac M4 Max, Snapdragon X — cut through the noise with rankings that respect your VRAM, memory, and target speed.

02 CALCULATOR

Cost ROI: self-hosted vs API

Sliders for your monthly token volume, electricity cost, GPU amortization. Real break-even point against GPT-5, Claude, Gemini, DeepSeek — updated pricing.

03 OPEN DATA

Public API & MCP server

178 JSON endpoints under CC BY 4.0, free to use in your own tools. Official MCP server on GitHub for ChatGPT, Claude Desktop, and Cursor.

04 METHOD

Independent benchmark pipeline

Continuous benchmarking against published model versions and quantizations. No press-kit numbers, no marketing decks — just tokens/sec backed by our open data API.

Learning center

79 guides, zero fluff.

Hands-on setup guides, hardware picks, and tool comparisons — filter by theme.

Showing 79 of 79 guides

Tools
Aider vs Continue.dev vs Cursor with Local LLMs (2026 Verdict)
Guides
Best Abliterated/Uncensored Local LLMs 2026
Guides
Best Free Local LLM in 2026
Rankings
Best LLM for Coding in 2026
Hardware
Best LLM for MacBook Air M3: 8GB, 16GB, or 24GB
Hardware
Best LLM for MacBook Pro M3 Pro and M3 Max (18GB-128GB)
Hardware
Best LLM for MacBook Pro M4 Pro and M4 Max (24GB–128GB)
Rankings
Best LLM for RAG in 2026
Hardware
Best LLM for RTX 5080 (16GB): What It Can Actually Run in 2026
Rankings
Best LLM for Translation in 2026
Rankings
Best LLM for Writing in 2026: The Definitive Verdict
Hardware
Best LLMs for Mac Mini M4 and M4 Pro (16–64GB) in 2026
Hardware
Best LLMs to Run on the RTX 4060 Ti (8GB vs 16GB)
Rankings
Best Local LLM for 16GB VRAM in 2026
Rankings
Best Local LLM for MacBook Pro M4 Max: Tested & Ranked
Rankings
Best Local LLM for RTX 3090 (24GB) in 2026
Rankings
Best Local LLM for RTX 4070 / 4070 Super / 4070 Ti (12GB)
Rankings
Best Local LLM for RTX 4090 (2026 Benchmarks)
Rankings
Best Local LLM for Smart Home & Privacy in 2026
Hardware
Best Local LLMs for the RTX 5070 Ti (16GB) in 2026
Hardware
Best MLX Local LLMs on Apple Silicon in 2026
Rankings
Best local LLM for GDPR compliance in 2026
Rankings
Best local LLM for RTX 5090: the 2026 verdict
Compliance
Can I Use This LLM Commercially? — License Decoder
Tools
Claude Desktop + Local LLM via MCP — Step-by-Step
Hardware
DeepSeek V4 Self-Hosted: VRAM Budget and Step-by-Step Deployment
Tools
ExLlama vs vLLM vs llama.cpp — Production Inference Compared
Guides
GLM 5.1 Local: Setup & Benchmarks
Tools
Gemma 4 Local: Ollama Setup & Benchmarks
Compliance
HIPAA-Compliant Local LLM Stack — What You Actually Need
Tools
How to Add a Local LLM to Claude Desktop via MCP
Tools
How to Build llama.cpp from Source with CUDA — 2026 Guide
Hardware
How to Build llama.cpp with Metal on Mac M-Series
Hardware
How to Install DeepSeek R1 32B on RTX 5090
Hardware
How to Install Llama 3.3 70B on RTX 4090 — Step-by-Step
Hardware
How to Install Mistral Small 3.1 24B on RTX 3060
Tools
How to Install Ollama on Windows 11 with WSL2 — Full Guide
Tools
How to Install Ollama with ROCm on AMD GPU
Hardware
How to Install Qwen 2.5 Coder 32B on Mac M4 Max
Guides
How to Install text-generation-webui with CUDA on Windows
Hardware
How to Pick a GPU for Local LLM in 2026 — Buyer's Guide
Hardware
How to Run Llama on Apple Silicon with MLX — Native Performance
Guides
How to Run a Local LLM: A 2026 Beginner Guide
Tools
How to Run vLLM in Docker with NVIDIA CUDA — 10-Minute Setup
Tools
How to Self-Host OpenWebUI in Docker — 5-Minute Setup
Tools
How to Self-Host TabbyAPI as an OpenAI-Compatible Endpoint
Tools
How to Use Aider with Ollama as a Local Copilot
Tools
How to Use LiteLLM Router with Ollama — OpenAI-Drop-In
Tools
How to Use LlamaIndex with a Local Ollama LLM
Tools
How to Use a Local Ollama LLM as Cursor's Backend
Tools
How to Wire Continue.dev to a Local Ollama LLM in VS Code
Tools
How to Wire LangChain to a Local LLM — Production-Ready
Guides
How to replace ChatGPT with a local LLM
Hardware
LLM Context Window vs VRAM Cost — 128k vs 32k Compared
Hardware
LLM VRAM Requirements — Every Model, Every Quant (2026)
Tools
LM Studio MTP: The Complete Multi-Token Prediction Guide
Tools
LlamaIndex vs LangChain with Local LLM — Which to Pick
Compliance
Local LLM for SOC 2 — What Your Auditor Wants to See
Compare
Local LLM vs ChatGPT: Which Should You Actually Run in 2026?
Tools
Ollama vs LM Studio vs llama.cpp — Honest 2026 Comparison
Tools
Ollama vs vLLM for Production Self-Hosting: The 2026 Verdict
Tools
OpenWebUI vs LibreChat — Best Self-Hosted ChatGPT Clone
Compare
Q4 vs Q5 vs Q8 Quantization — Real Quality Loss Tested
Guides
Qwen3.6 35B-A3B Local: Review & VRAM Requirements
Guides
Qwen3.7 Max: Best Open-Weight to Self-Host
Tools
SillyTavern + Local LLM — Best Setup for Role-Play
Tools
TabbyAPI — Self-Host an OpenAI-Compatible API in 10 Minutes
Compare
TurboQuant Quantization Explained: Run Big Models on Less VRAM
Tools
Use LiteLLM Router with Your Local LLM — Drop-In OpenAI Replacement
Tools
Use Your Own Local LLM in Cursor — BYO Model Setup
Tools
We Switched From Ollama to llama.cpp — What Broke
Hardware
What Actually Fits in 12GB of VRAM: Measured Numbers from a 5070 Ti Laptop
Hardware
What LLM Can You Run on 8GB VRAM in 2026?
Hardware
What LLMs Can You Run on the RTX 3060 12GB in 2026?
Hardware
Which LLM Runs Best on the MacBook Air M4 (16GB, 24GB, or 32GB)?
Guides
Which LLM Runs Best on the Radeon RX 7900 XTX (24GB)?
Hardware
Which LLM Should You Run on a Mac Studio (M2/M3/M4 Ultra)?
Hardware
Which LLM Should You Run on an RTX 4080 or 4080 Super?
Hardware
Why VRAM Matters More Than TFLOPS for LLM Inference
All guides →
The catalog

239 models, every angle.

The catalog's most-tracked families — one flagship model per author. Filter and jump straight into the full catalog.

397B · Apache 2.0
Qwen 3.5 397B-A17B
240 GB Q4 · 255k ctx
31B · Gemma
Gemma 4 31B
18 GB Q4 · 250k ctx
561B · NVIDIA Open Model License
Nemotron 3 Ultra Base (BF16)
325 GB Q4 · 125k ctx
1700B · MIT
DeepSeek V4 Pro 0813 1.7T
986 GB Q4 · 1024k ctx
32B · Apache 2.0
Granite 4.0 H-Small 32B-A9B
19 GB Q4 · 125k ctx
675B · Apache 2.0
Mistral Large 3 675B
405 GB Q4 · 250k ctx
405B · Llama 3.1 Community
Llama 3.1 405B Instruct
240 GB Q4 · 125k ctx
72B · Apache 2.0
Molmo 72B
42 GB Q4 · 4k ctx
24B · LFM Open License v1.0
LFM2 24B
14 GB Q4 · 32k ctx
14B · MIT
Phi-4 Reasoning 14B
9 GB Q4 · 32k ctx
35B · CC-BY-NC 4.0
Aya 23 35B
20 GB Q4 · 8k ctx
753B · MIT
GLM 5.2 753B-A40B
437 GB Q4 · 976k ctx
2800B · Kimi License
Kimi K3
1624 GB Q4 · 976k ctx
8B · MiniCPM Model License
MiniCPM-o 2.6 8B
5.5 GB Q4 · 31k ctx
10B · TII Falcon-LLM License 2.0
Falcon 3 10B Instruct
6 GB Q4 · 31k ctx
406B · Tencent Hunyuan License
Hunyuan Large 2.0
245 GB Q4 · 256k ctx
104B · CC-BY-NC 4.0
Command R+ 104B (08-2024)
60 GB Q4 · 125k ctx
3B · Apache 2.0
SmolLM3 3B
2 GB Q4 · 125k ctx
118B · OpenMDW 1.1
Laguna S 2.1
68 GB Q4 · 256k ctx
1020B · MIT
MiMo V2.5 Pro
595 GB Q4 · 976k ctx
34B · Apache 2.0
Yi 1.5 34B Chat
20 GB Q4 · 4k ctx
1000B · MIT
Ling 2.6 1T
580 GB Q4 · 256k ctx
40B · Apache 2.0
Salamandra 40B Instruct
24 GB Q4 · 8k ctx
300B · Apache 2.0
ERNIE 4.5 300B-A47B
180 GB Q4 · 128k ctx
Browse all 239 models →
Who's behind this

Independent. Skin in the game.

BestLLMfor is built and operated by Mohamed Meguedmi — one engineer, a continuous benchmark pipeline, a public data API and an open-source MCP server.

No VC, no SEO farm. One engineer obsessed with tracking every model worth running, and publishing what the numbers say — transparently.

BENCHMARK PIPELINE
Models tracked239+ (daily)
Quants testedQ4 · Q5 · Q8 · FP16
Data API178 JSON · CC BY 4.0
MCP serverPublic · open source
Methodologysee how →
Newsletter

Your GPU cheat sheet,
then a hands-on series.

Get the VRAM-to-model cheat sheet by email, then a short series to build your local copilot — plus occasional model drops & benchmarks. Unsubscribe anytime, one click.

Want the deep dive instead? See the Local Copilot Kit →