The market at a glance
All the news →The QuelLLM kits — the reference guide per use case.
One kit for each use case: the complete guide, ready-to-paste configurations, and lifetime online space.
The catalog of open-weight LLMs
that run locally.
All relevant models, with required VRAM by quantization, estimated speed, and use cases. French models are featured.
Find your LLM by your needs
Don't want to use the configurator? Our topic-based rankings select the best self-hostable models by use case and hardware.
Documentation QuelLLM.
341+ tutorials for Windows, macOS, and Linux. The tests performed are specified in each guide. From the initial installation to advanced RAG and fine-tuning techniques.
Featured
— our essential guidesInstall Ollama on Windows 11: complete guide (2026)
Download Ollama for Windows 11, install it, enable the GPU (CUDA), and run your first model: the 2026 step-by-step guide, including common errors.
Which LLM on RTX 4090 (24 GB)?
RTX 4090 24 GB, the leading mainstream local AI option in 2026. Qwen 3.8 27B, Devstral 24B, gpt-oss 20B, GLM 4.7 Flash, tokens/sec benchmarks, CUDA Ada optimizations.
Which LLM on RTX 5090 (32 GB)?
RTX 5090 2025: 32 GB of GDDR7 VRAM, 1792 GB/s bandwidth. Run large 2026 MoEs (Qwen 3.6 35B-A3B, Nemotron 3.5 Lightning) and Qwen 3.8 27B with long context locally, benchmarks, ideal configurations.
DeepSeek V4 Pro locally: VRAM, hardware, installation
What machine do you need for DeepSeek V4 Pro (1.6T MoE, 49B active, MIT)? Required VRAM/RAM, local installation, benchmarks vs. GPT-5, and real-world limitations.
LM Studio for beginners: your first local chat in 10 minutes
Your first local LLM with LM Studio: install it, download a model, and chat in 10 minutes. No command line, ideal for beginners.
Which LLM for RTX 3060 with 12 GB?
RTX 3060 12 GB: the iconic budget LLM GPU. 12 GB for €250 used, Gemma 4 12B, Qwen 3.5 9B, Granite 4.2 8B, detailed benchmarks.
DeepSeek V4 Flash 284B : the 1st frontier that fits on Mac Studio
DeepSeek V4 Flash 284B MoE (13B active, MIT, 1M ctx): the first frontier model executable on a workstation. Mac Studio Ultra installation, benchmarks, Pro comparison.
Install Ollama on macOS (Apple Silicon): 2026 guide
Leverage Metal and M1/M2/M3/M4 unified memory.
Which LLM on Mac mini M4 / M4 Pro (16–64 GB)?
Mac mini M4: the best price-to-performance ratio for local AI in 2026. Benchmarks, recommended configuration, home server use.
Which LLM on RTX 3090 / 3090 Ti (24 GB)?
RTX 3090 and used 3090 Ti 24 GB: still excellent for LLMs in 2026. Qwen 3.8 27B, Qwen3-Coder 30B, benchmarks, performance/price verdict, cooling.
Which LLM for 12 GB of VRAM?
12 GB VRAM (RTX 3060 12GB, 4070, 5070): the 2026 sweet spot. Qwen 3.5 9B in Q8, Gemma 4 12B, multi-stage RAG. The definitive guide.
Qwen 3.8 27B locally: Ollama installation and VRAM
Run Qwen3.8-27B locally: exact Ollama command, VRAM by quantization, MLX variant on Mac, the 256k context trap, thinking mode, and official benchmarks.
Your first local conversation
Launch Ollama, load Mistral, and start chatting. The day 1 tutorial.
Which LLM on RTX 5070 Ti (16 GB)?
RTX 5070 Ti 16 GB: the 2025 sweet spot for local AI. Ollama benchmarks, comfortable 14B/24B models, comparison with 4070 Ti Super.
Which LLM on MacBook Pro M4 Pro / Max (24–128 GB)?
MacBook Pro M4 Pro / Max 2025: 546 GB/s bandwidth, which models really leverage the chip, practical limitations.
Which LLM on RTX 4070 / 4070 Super / 4070 Ti (12 GB)?
RTX 4070, 4070 Super, and 4070 Ti 12 GB: LLM comparison, 13B models that run comfortably, 12 GB limitations, measured benchmarks.
Ollama vs. LM Studio vs. Jan vs. GPT4All
Summary table for choosing the right tool for your profile.
Which LLM for 8 GB of VRAM?
The complete guide to 8 GB of VRAM (RTX 3050/3060 8GB, 4060, 5050, 5060): Qwen 3.5 9B, Granite 4.2 8B, tips for stretching VRAM.
Enterprise coding AI: protect proprietary code, NDAs, and trade secrets
Deploy 100% local coding AI (Ollama + Cline + Aider + Tabby) for a development team under an NDA and trade-secret protection: why Copilot and Cursor expose your proprietary code, what the GDPR and AI Act require, the Tabby workstation-vs-GPU-server architecture, and the non-exfiltration audit.
Install Ollama in 5 minutes (Windows, macOS, Linux)
How to install Ollama in 5 minutes on Windows, macOS, and Linux: RAM/GPU requirements, essential commands, and first local model to run right afterward.
Prompting basics
Structure your prompts to get useful answers.
Which LLM for MacBook Pro M3 Pro / Max (18–128 GB)?
MacBook Pro M3 Pro / Max: the best laptop for local AI in 2026. Qwen 3.6 35B-A3B, 128k context, Flash Attention.
Which LLM on RTX 5080 (16 GB)?
RTX 5080 Blackwell: 16 GB GDDR7 at 960 GB/s. Qwen 3.5, Granite 4.2, Gemma 4, gpt-oss, Devstral benchmarks in Q4. Optimal Ollama configuration.
Which LLM for 16 GB of VRAM?
16 GB VRAM (RTX 4070 Ti Super, 5070 Ti, 5080, 4060 Ti 16GB): Mistral Small 24B, Devstral 24B, gpt-oss 20B, the 2026 pro tier.
Which GPU for a local LLM? RTX 4070 to 5090, Mac (2026)
RTX 4070 vs 4090 vs Mac M-Max: the 2026 buying guide.
Local RAG: introduction
Understand Retrieval-Augmented Generation to chat with your docs.
Which LLM on Mac Studio (M2 / M3 / M4 Ultra, 64–512 GB)?
Mac Studio Ultra: up to 512 GB of unified memory. Run Llama 4 Maverick 402B, Qwen3-235B, DeepSeek 671B locally. The power user's guide.
Which LLM on RTX 4080 / 4080 Super (16 GB)?
RTX 4080 and 4080 Super 16 GB for local LLMs: every model that fits, benchmarks, 4080 vs. 4080 Super comparison, buying verdict.
Which LLM on Radeon RX 7900 XTX (24 GB)?
Radeon RX 7900 XTX 24 GB: AMD alternative to RTX 4090 for LLMs. ROCm 6.x, Qwen 3.8 27B, tokens/sec benchmarks, 2026 verdict.
Which LLM for 24 GB of VRAM?
24 GB VRAM (RTX 3090, 4090, RX 7900 XTX): Qwen 3.8 27B entirely in VRAM, large 35B-A3B MoE, LoRA fine-tuning. The serious tier.
Reviews & tests.
When local deployment isn't enough, which online services are worth your money? 14 brands reviewed, 7 rejected. Each profile leads with the free local alternative.
NVIDIA GB10 chip and 128 GB of unified memory: what this machine can really run, measured on our unit with public data to back it up. Machine generously provided by GIGABYTE; the measurements and opinions are ours.
The full AI TOP ATOM spec sheet →- Oct 81. Which LLMs really run on 128 GBRead →
- Oct 152. AI TOP ATOM and DGX SparkComing soon
- Oct 223. GB10, Ryzen AI Max+ 395, or MacComing soon
- Oct 294. A 70B LLM running locallyComing soon
- Nov 55. Local fine-tuningComing soon
- Nov 126. Private AI assistant for SMBsComing soon
- Nov 197. Local AI power consumptionComing soon
Three paths,
depending on who you are.
Each path is a sequence of guides designed for a specific profile. From the first download to an operational setup.
« I just want to try it, hassle-free. »
You've heard about local LLMs and want to see how they perform on your machine. No code, no system configuration—in 10 minutes, you'll be chatting with your first model.