Tool

Face-to-face models

Choose two or three models, and compare them using the same criteria. Parameters, VRAM by quantization, license, languages—everything that matters for making a choice.

Model A
Mistral 7B Instruct
Mistral AI · 7B
Model B
Gemma 3 27B
Google · 27B
Model C
Llama 3.3 70B Instruct
Meta · 70B
A
Mistral 7B Instruct
B
Gemma 3 27B
C
Llama 3.3 70B Instruct
Family
Mistral
Gemma
Llama
Editor
Mistral AI
Google
Meta
Origin
🇫🇷 FR
🇺🇸 US
🇺🇸 US
License
Apache 2.0
Gemma
Llama 3.3 Community
Settings
7B
27B
70B
Context
33k tokens
128k tokens
128k tokens
VRAM (Q4)
5 GB
16 GB
40 GB
VRAM (Q5)
6 GB
19 GB
48 GB
VRAM (Q8)
9 GB
29 GB
75 GB
VRAM (FP16)
16 GB
54 GB
140 GB
Minimum CPU RAM
8 GB
28 GB
64 GB
Tok/s (average)
35 tok/s
13 tok/s
6 tok/s
Tags
chat, general
chat, general, vision, multilingual
chat, general, reasoning
Prewritten comparisons

Duels in one click

Instead of reconfiguring the tool every time, consult our comparisons, already written up—with verdicts by use case, VRAM, tokens/sec, and FAQ.

Summary verdict
The smallest
Mistral 7B Instruct
The classic French model. Fast, versatile, and an excellent base to get started.
The all-rounder
Gemma 3 27B
High-end Gemma. LMArena Elo 1338 — beats Llama 3.1 405B at 15× smaller size.
The powerhouse
Llama 3.3 70B Instruct
Llama 3.1 405B quality at 1/6 the size. Weights under a community license, gated HF access.