All models · 100% local, 100% private

Which LLM runs on
your machine

Tell us what's under the hood. We'll tell you what runs, how fast, and how to install it—step by step.

Don't have the machine yet? Which PC for local AI, based on your budget →
128 GB mini PC · NVIDIA DGX Spark · All AI hardware

~/configurateur —— ● ready
Computation stays in your browser. The site also uses audience measurement.
Free
No sign-up
In English
Educational guides
Open source
Public data
Local compute
Audience measurement
◆ QuelLLM kits — the definitive guide by use case · $24 per kit · $49 all for life →

The market at a glance

All the news →
9 oct.TII launches Falcon ASR, an open-weight speech recognition model9 oct.Nvidia-tuned Nemotron earns two gold-level results at the IOI and IMO9 oct.Liquid AI releases open d1, open multimodal decision models for the edge
Catalog · 253 models

The catalog of open-weight LLMs
that run locally.

All relevant models, with required VRAM by quantization, estimated speed, and use cases. French models are featured.

🇫🇷
Sovereignty focus

The 27 models made in France

Mistral, Lucie (OpenLLM-France), CroissantLLM — trained on French-language corpora, permissive licenses, excellent for formal French.

★Top editorial picksView all rankings →
253 models
Model↕
Params↕
Q4 VRAM↕
Context↕
Runs on
tok/s↕
Output↕
Pleias-RAG 1B★ FRchatfr
🇫🇷 PleIAs · Specialized 1.2B RAG model. Citation and grounding built in. Outperforms most SLMs ≤4B on HotPotQA.
1.2B
0.8 GB
2k
8121624
150
avr. 2025
CroissantLLM 1.3B★ FRchatfr
🇫🇷 CroissantLLM · Small bilingual FR/EN model. Runs anywhere, even on a CPU.
1.3B
1 GB
2k
8121624
120
janv. 2024
SmolLM2 1.7B Instruct★ FRchatgeneral
🇫🇷 HuggingFace · Highly downloaded 1.7B Apache 2.0 model. Beats Qwen2.5-1.5B by ~6 MMLU-Pro points. BFCL function calling 27%.
1.7B
1.2 GB
8k
8121624
120
nov. 2024
Helium 1 2B★ FRchatgeneral
🇫🇷 Kyutai · Multilingual base for 24 EU languages (Kyutai FR). Distilled from Gemma 2 → Gemma Terms also apply.
2B
1.5 GB
4k
8121624
100
avr. 2025
SmolVLM2 2.2B Instruct★ FRvisionchat
🇫🇷 HuggingFace · 2.2B VLM: image + video + text. 5.2 GB VRAM for video inference. SmolLM2-1.7B base.
2.2B
1.6 GB
8k
8121624
90
févr. 2025
Pleias 3B Preview★ FRchatmultilingual
🇫🇷 PleIAs · French. Trained 100% on Common Corpus (open data). EU AI Act compliant by design.
3B
2 GB
2k
8121624
70
déc. 2024
SmolLM3 3B★ FRchatgeneral
🇫🇷 HuggingFace · 3B dual-mode (think/no-think). 6 languages. MMLU 59.7, GSM8K 70.9. Fully open (data + recipe).
3B
2 GB
125k
8121624
70
juil. 2025
Shieldstral 3B★ FRgeneralsmall
🇫🇷 Mistral AI · Mistral 3B guardrail model (content moderation), 32k context, ~1.7 GB Q4 VRAM. Self-hosted, Apache 2.0 license.
3B
1.7 GB
32k
8121624
85
—
Voxtral-4B-TTS★ FRaudiomultilingual
🇫🇷 Mistral AI · Open-frontier TTS, 9 languages including FR. Rivals ElevenLabs. ⚠ Non-commercial.
4B
3 GB
4k
8121624
65
mars 2026
Mistral 7B Instruct★ FRchatgeneral
🇫🇷 Mistral AI · The classic French model. Fast, versatile, and an excellent base to get started.
7B
5 GB
32k
8121624
35
sept. 2023
Lucie 7B★ FRchatfr
🇫🇷 OpenLLM-France · Sovereign French-language LLM, trained on French corpora.
7B
5 GB
4k
8121624
35
janv. 2025
Codestral Mamba 7B★ FRcodefr
🇫🇷 Mistral AI · Pure Mamba SSM for code. Linear inference, 256k ctx. No Ollama (partial llama.cpp support).
7B
5 GB
250k
8121624
40
juil. 2024
Claire 7B 0.1★ FRchatfr
🇫🇷 LINAGORA · LoRA-finetuned Falcon-7B on spontaneous French dialogue. ⚠ NC-SA license. Separate Apache variant.
7B
5 GB
2k
8121624
35
nov. 2023
Moshi 7B★ FRaudiofr
🇫🇷 Kyutai · First open full-duplex speech model. ~200 ms latency. Moshiko/Moshika voices. Kyutai (French lab).
7.6B
5 GB
4k
8121624
25
sept. 2024
Mistral Nemo 12B Instruct★ FRchatgeneral
🇫🇷 Mistral AI · Co-developed with NVIDIA. 128k ctx, Tekken tokenizer, strong in European multilingual.
12B
7 GB
125k
8121624
25
juil. 2024
Codestral 22B v0.1★ FRcodefr
🇫🇷 Mistral AI · Code 22B Mistral, 80+ languages. ⚠ MNPL non-production license — personal/research use.
22B
13 GB
31.25k
8121624
16
mai 2024
Mistral Small 3★ FRchatgeneral
🇫🇷 Mistral AI · Excellent quality-to-size ratio at release (early 2025). Competes with 70B models.
24B
14 GB
32k
8121624
15
janv. 2025
Mistral Small 3.1 24B★ FRchatgeneral
🇫🇷 Mistral AI · Small 3 enhanced with vision. 128k ctx, Apache 2.0. Replaced by Small 3.2 since June 2025.
24B
14 GB
125k
8121624
15
mars 2025
Devstral Small 2 24B★ FRcodefr
🇫🇷 Mistral AI · 24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.
24B
14 GB
250k
8121624
15
déc. 2025
Mistral Small 3.2 24B★ FRchatgeneral
🇫🇷 Mistral AI · June 2025 update for Small 3.1. Half as many infinite generations, improved function calling.
24B
14 GB
125k
8121624
15
juin 2025
Magistral Small 24B★ FRreasoningfr
🇫🇷 Mistral AI · First open Mistral reasoner. AIME24 70.7%. Based on Small 3.1 + CoT training.
24B
14 GB
125k
8121624
15
juin 2025
Mixtral 8x7B★ FRchatgeneral
🇫🇷 Mistral AI · 8-expert MoE. High quality, but VRAM-hungry.
47B
26 GB
32k
8121624
12
déc. 2023
Mistral Small 4★ FRchatgeneral
🇫🇷 Mistral AI · 119B/6.5B active MoE unifies chat + reasoning + vision + code. The French flagship of 2026.
119B
72 GB
250k
8121624
12
mars 2026
Mistral Medium 3.5 128B★ FRchatgeneral
🇫🇷 Mistral AI · Dense 128B + vision, 256k ctx, configurable reasoning. SWE-Bench 77.6%. Replaces Medium 3.1 and Magistral. Released April 29, 2026.
128B
74 GB
250k
8121624
4
avr. 2026
Mixtral 8x22B Instruct★ FRchatgeneral
🇫🇷 Mistral AI · Apache 2.0 MoE, 141B/39B active. MMLU 77.8, HumanEval 45.1. 80 GB in Q4.
141B
82 GB
62.5k
8121624
8
avr. 2024
Mistral Large 3 675B★ FRchatgeneral
🇫🇷 Mistral AI · 675B/41B active MoE + 2.5B vision encoder, Apache 2.0. #2 OSS non-reasoning model on LMArena. Trained on 3,000 H200s.
675B
405 GB
250k
8121624
5
déc. 2025
Mistral Large 4★ FRchatgeneral
🇫🇷 Mistral AI · 1T/49B-active MoE from Mistral AI, 32k context, ~580 GB Q4 VRAM: frontier flagship reserved for multi-GPU servers.
1000B
580 GB
32k
8121624
10
oct. 2026
Granite Time Series PatchTST R2generalsmall
🇺🇸 IBM · IBM foundation model for time series (~385M, PatchTST), multivariate forecasting, ~0.2 GB VRAM Q4. Ultra-lightweight, runs on CPU. Apache 2.0 license.
0.4B
0.2 GB
8k
8121624
170
août 2026
Activity Generation 0.5B (Qwen)chatgeneral
🇨🇳 mjpsm · Fine-tune Qwen 0.5B specialized for activity generation, 32k context, ~0.3 GB Q4 VRAM. Ultra-lightweight, runs even on CPU. Apache 2.0 license.
0.5B
0.3 GB
32k
8121624
170
sept. 2026
S1-miniaudiosmall
🇺🇸 superwhisper · S1-mini (superwhisper): 0.6B speech recognition based on Qwen3-0.6B, 40k context, ~0.3 GB VRAM Q4. Ultra-lightweight local transcription.
0.6B
0.3 GB
40k
8121624
170
août 2026
Learning center

Documentation QuelLLM.

341+ tutorials for Windows, macOS, and Linux. The tests performed are specified in each guide. From the initial installation to advanced RAG and fine-tuning techniques.

⌘ K

Featured

— our essential guides
★ Featured
Enterprise · Deployment

Enterprise coding AI: protect proprietary code, NDAs, and trade secrets

Deploy 100% local coding AI (Ollama + Cline + Aider + Tabby) for a development team under an NDA and trade-secret protection: why Copilot and Cursor expose your proprietary code, what the GDPR and AI Act require, the Tabby workstation-vs-GPU-server architecture, and the non-exfiltration audit.

Advanced·16 min·Mohamed Meguedmi
★ Featured
Getting started · Free

Best free local AI: top picks for 2026, even without a GPU

Local AI: definition and the best free options in 2026, to install on PC or Mac, even without a dedicated GPU. Tool comparison and step-by-step guide.

Beginner·14 min·Mohamed Meguedmi
★ Featured
Installation · Ollama

Install Ollama on Windows 11: complete guide (2026)

Download Ollama for Windows 11, install it, enable the GPU (CUDA), and run your first model: the 2026 step-by-step guide, including common errors.

Beginner·3 min·Mohamed Meguedmi
341 results
Beginner 3 min

Install Ollama on Windows 11: complete guide (2026)

Download Ollama for Windows 11, install it, enable the GPU (CUDA), and run your first model: the 2026 step-by-step guide, including common errors.

OllamaRead →
Intermediate 12 min

Which LLM on RTX 4090 (24 GB)?

RTX 4090 24 GB, the leading mainstream local AI option in 2026. Qwen 3.8 27B, Devstral 24B, gpt-oss 20B, GLM 4.7 Flash, tokens/sec benchmarks, CUDA Ada optimizations.

RTX 40Read →
Intermediate 11 min

Which LLM on RTX 5090 (32 GB)?

RTX 5090 2025: 32 GB of GDDR7 VRAM, 1792 GB/s bandwidth. Run large 2026 MoEs (Qwen 3.6 35B-A3B, Nemotron 3.5 Lightning) and Qwen 3.8 27B with long context locally, benchmarks, ideal configurations.

RTX 50Read →
Advanced 16 min

DeepSeek V4 Pro locally: VRAM, hardware, installation

What machine do you need for DeepSeek V4 Pro (1.6T MoE, 49B active, MIT)? Required VRAM/RAM, local installation, benchmarks vs. GPT-5, and real-world limitations.

DeepSeekRead →
Beginner 5 min

LM Studio for beginners: your first local chat in 10 minutes

Your first local LLM with LM Studio: install it, download a model, and chat in 10 minutes. No command line, ideal for beginners.

LM StudioRead →
Beginner 14 min

Which LLM for RTX 3060 with 12 GB?

RTX 3060 12 GB: the iconic budget LLM GPU. 12 GB for €250 used, Gemma 4 12B, Qwen 3.5 9B, Granite 4.2 8B, detailed benchmarks.

RTX 30Read →
Intermediate 14 min

DeepSeek V4 Flash 284B : the 1st frontier that fits on Mac Studio

DeepSeek V4 Flash 284B MoE (13B active, MIT, 1M ctx): the first frontier model executable on a workstation. Mac Studio Ultra installation, benchmarks, Pro comparison.

DeepSeekRead →
Beginner 3 min

Install Ollama on macOS (Apple Silicon): 2026 guide

Leverage Metal and M1/M2/M3/M4 unified memory.

OllamaRead →
Beginner 14 min

Which LLM on Mac mini M4 / M4 Pro (16–64 GB)?

Mac mini M4: the best price-to-performance ratio for local AI in 2026. Benchmarks, recommended configuration, home server use.

Mac miniRead →
Intermediate 13 min

Which LLM on RTX 3090 / 3090 Ti (24 GB)?

RTX 3090 and used 3090 Ti 24 GB: still excellent for LLMs in 2026. Qwen 3.8 27B, Qwen3-Coder 30B, benchmarks, performance/price verdict, cooling.

RTX 30Read →
Beginner 11 min

Which LLM for 12 GB of VRAM?

12 GB VRAM (RTX 3060 12GB, 4070, 5070): the 2026 sweet spot. Qwen 3.5 9B in Q8, Gemma 4 12B, multi-stage RAG. The definitive guide.

By VRAMRead →
Intermediate 14 min

Qwen 3.8 27B locally: Ollama installation and VRAM

Run Qwen3.8-27B locally: exact Ollama command, VRAM by quantization, MLX variant on Mac, the 256k context trap, thinking mode, and official benchmarks.

OllamaRead →
Beginner 10 min

Your first local conversation

Launch Ollama, load Mistral, and start chatting. The day 1 tutorial.

PromptingRead →
Intermediate 11 min

Which LLM on RTX 5070 Ti (16 GB)?

RTX 5070 Ti 16 GB: the 2025 sweet spot for local AI. Ollama benchmarks, comfortable 14B/24B models, comparison with 4070 Ti Super.

RTX 50Read →
Intermediate 11 min

Which LLM on MacBook Pro M4 Pro / Max (24–128 GB)?

MacBook Pro M4 Pro / Max 2025: 546 GB/s bandwidth, which models really leverage the chip, practical limitations.

MacBook ProRead →
Intermediate 14 min

Which LLM on RTX 4070 / 4070 Super / 4070 Ti (12 GB)?

RTX 4070, 4070 Super, and 4070 Ti 12 GB: LLM comparison, 13B models that run comfortably, 12 GB limitations, measured benchmarks.

RTX 40Read →
Beginner 11 min

Ollama vs. LM Studio vs. Jan vs. GPT4All

Summary table for choosing the right tool for your profile.

ToolsRead →
Beginner 11 min

Which LLM for 8 GB of VRAM?

The complete guide to 8 GB of VRAM (RTX 3050/3060 8GB, 4060, 5050, 5060): Qwen 3.5 9B, Granite 4.2 8B, tips for stretching VRAM.

By VRAMRead →
Advanced 16 min

Enterprise coding AI: protect proprietary code, NDAs, and trade secrets

Deploy 100% local coding AI (Ollama + Cline + Aider + Tabby) for a development team under an NDA and trade-secret protection: why Copilot and Cursor expose your proprietary code, what the GDPR and AI Act require, the Tabby workstation-vs-GPU-server architecture, and the non-exfiltration audit.

DeploymentRead →
Beginner 11 min

Install Ollama in 5 minutes (Windows, macOS, Linux)

How to install Ollama in 5 minutes on Windows, macOS, and Linux: RAM/GPU requirements, essential commands, and first local model to run right afterward.

OllamaRead →
Beginner 11 min

Prompting basics

Structure your prompts to get useful answers.

PromptingRead →
Intermediate 12 min

Which LLM for MacBook Pro M3 Pro / Max (18–128 GB)?

MacBook Pro M3 Pro / Max: the best laptop for local AI in 2026. Qwen 3.6 35B-A3B, 128k context, Flash Attention.

MacBook ProRead →
Intermediate 11 min

Which LLM on RTX 5080 (16 GB)?

RTX 5080 Blackwell: 16 GB GDDR7 at 960 GB/s. Qwen 3.5, Granite 4.2, Gemma 4, gpt-oss, Devstral benchmarks in Q4. Optimal Ollama configuration.

RTX 50Read →
Intermediate 11 min

Which LLM for 16 GB of VRAM?

16 GB VRAM (RTX 4070 Ti Super, 5070 Ti, 5080, 4060 Ti 16GB): Mistral Small 24B, Devstral 24B, gpt-oss 20B, the 2026 pro tier.

By VRAMRead →
Beginner 12 min

Which GPU for a local LLM? RTX 4070 to 5090, Mac (2026)

RTX 4070 vs 4090 vs Mac M-Max: the 2026 buying guide.

GPURead →
Beginner 11 min

Local RAG: introduction

Understand Retrieval-Augmented Generation to chat with your docs.

ConceptsRead →
Advanced 13 min

Which LLM on Mac Studio (M2 / M3 / M4 Ultra, 64–512 GB)?

Mac Studio Ultra: up to 512 GB of unified memory. Run Llama 4 Maverick 402B, Qwen3-235B, DeepSeek 671B locally. The power user's guide.

Mac StudioRead →
Intermediate 11 min

Which LLM on RTX 4080 / 4080 Super (16 GB)?

RTX 4080 and 4080 Super 16 GB for local LLMs: every model that fits, benchmarks, 4080 vs. 4080 Super comparison, buying verdict.

RTX 40Read →
Intermediate 11 min

Which LLM on Radeon RX 7900 XTX (24 GB)?

Radeon RX 7900 XTX 24 GB: AMD alternative to RTX 4090 for LLMs. ROCm 6.x, Qwen 3.8 27B, tokens/sec benchmarks, 2026 verdict.

Radeon RX 7000Read →
Intermediate 11 min

Which LLM for 24 GB of VRAM?

24 GB VRAM (RTX 3090, 4090, RX 7900 XTX): Qwen 3.8 27B entirely in VRAM, large 35B-A3B MoE, LoRA fine-tuning. The serious tier.

By VRAMRead →
…
1–30 on 341 guides
Test bench

Reviews & tests.

When local deployment isn't enough, which online services are worth your money? 14 brands reviewed, 7 rejected. Each profile leads with the free local alternative.

Hardware7 guides · one new guide every Thursday
GIGABYTE AI TOP ATOM: our test bench

NVIDIA GB10 chip and 128 GB of unified memory: what this machine can really run, measured on our unit with public data to back it up. Machine generously provided by GIGABYTE; the measurements and opinions are ours.

The full AI TOP ATOM spec sheet →
  1. 8 oct.1. Quels LLM tournent vraiment sur 128 GoRead →
  2. 15 oct.2. AI TOP ATOM et DGX SparkComing soon
  3. 22 oct.3. GB10, Ryzen AI Max+ 395 ou MacComing soon
  4. 29 oct.4. Un LLM 70B en localComing soon
  5. 5 nov.5. Fine-tuning en localComing soon
  6. 12 nov.6. Assistant IA privé pour PMEComing soon
  7. 19 nov.7. Consommation d’une IA localeComing soon
API & code4/5
GLM Coding Plan
The API for GLM open-weight models, tested on our hardware.
Transcription4.5/5
HappyScribe
Cloud transcription versus local Whisper: our data-driven comparison.
AI agents3.5/5
Creao AI
Agents connected to your tools, no coding required. Still early.
Training3.5/5
Code Labs Academy
Accredited data/AI bootcamp for pursuing a career change.
See the methodology and all reviews →
Guided paths

Three paths,
depending on who you are.

Each path is a sequence of guides designed for a specific profile. From the first download to an operational setup.

01
Pathways curious

« I just want to try it, hassle-free. »

You've heard about local LLMs and want to see how they perform on your machine. No code, no system configuration—in 10 minutes, you'll be chatting with your first model.

Total duration
~15 min
Prerequisites
None
What you'll have at the end
Ollama installed
Mistral 7B launched
First successful prompt
Start the journey→