Which LLM runs on
your machine?
Tell us what's under the hood. We'll tell you what runs, how fast, and how to install it — step by step, in plain English.
The market in brief
All briefs →Built around your decision, not vendor benchmarks.
Four practical tools that answer the questions you actually have when picking an LLM.
Hardware-matched rankings
Best local LLM for RTX 4090, RTX 5090, Mac M4 Max, Snapdragon X — cut through the noise with rankings that respect your VRAM, memory, and target speed.
Cost ROI: self-hosted vs API
Sliders for your monthly token volume, electricity cost, GPU amortization. Real break-even point against GPT-5, Claude, Gemini, DeepSeek — updated pricing.
Public API & MCP server
178 JSON endpoints under CC BY 4.0, free to use in your own tools. Official MCP server on GitHub for ChatGPT, Claude Desktop, and Cursor.
Independent benchmark pipeline
Continuous benchmarking against published model versions and quantizations. No press-kit numbers, no marketing decks — just tokens/sec backed by our open data API.
79 guides, zero fluff.
Hands-on setup guides, hardware picks, and tool comparisons — filter by theme.
Showing 79 of 79 guides
239 models, every angle.
The catalog's most-tracked families — one flagship model per author. Filter and jump straight into the full catalog.
Pick your path.
Three common starting points — jump straight to the one that matches you.
New to local LLMs?
Start here: what "local" means, what hardware you need, and your first model in 10 minutes.
Coding & dev work
Ranked local models for autocomplete, agents, and full coding sessions — matched to your GPU.
Replace ChatGPT at work
Keep sensitive data in-house. What actually works as a private, self-hosted swap-in.
Independent. Skin in the game.
BestLLMfor is built and operated by Mohamed Meguedmi — one engineer, a continuous benchmark pipeline, a public data API and an open-source MCP server.
No VC, no SEO farm. One engineer obsessed with tracking every model worth running, and publishing what the numbers say — transparently.
| Models tracked | 239+ (daily) |
| Quants tested | Q4 · Q5 · Q8 · FP16 |
| Data API | 178 JSON · CC BY 4.0 |
| MCP server | Public · open source |
| Methodology | see how → |
Your GPU cheat sheet,
then a hands-on series.
Get the VRAM-to-model cheat sheet by email, then a short series to build your local copilot — plus occasional model drops & benchmarks. Unsubscribe anytime, one click.
Want the deep dive instead? See the Local Copilot Kit →