How we measure

Some numbers verifiable

This site displays figures: configurator estimates, catalog verdicts, and measured benchmarks. Here’s exactly where they come from and how we calculate them.

See the measured driver and its limits · Reproducible protocol · Public engine

AI TOP ATOM series (GIGABYTE), October 2026

For our series of 7 guides on AI TOP ATOM, provided free of charge by GIGABYTE and published at a rate of one guide per week since October 8, 2026, the protocol was finalized on October 6, 2026, at 13:30, before the first measurement. The public methodology reproduces its rules, the methods used in each guide, and the dated list of amendments. GIGABYTE neither reviewed nor approved our guides before publication. Our data, infographics, and photos from the series are licensed CC BY 4.0 : free reuse with attribution to quelllm.fr.

Measurement method and dated addenda · SHA-256 fingerprints of the models · Table data (CSV) · AI TOP ATOM datasheet

Three principles

  1. 01

    Estimated and measured values are distinguished

    The configurator provides tokens/sec estimated by formula. The Benchmarks page shows measured tokens/sec. They’re labeled differently so you don’t confuse them.

  2. 02

    We display the margin of error

    We do not have a calibrated error interval for all estimates. The measured driver reports the median, minimum, and maximum of the five repetitions.

  3. 03

    We date everything

    Every guide and every model page includes a date for the last verification. What was true 6 months ago about Ollama is often no longer true today.

The formulas used

Rules for the public compatibility engine. They assess memory and indicative speed; they are not a quality ranking.

01
Model weights memory budget
budget ≥ the catalog threshold for the quantization

The Q4/Q5/Q8/FP16 thresholds are catalog approximations. The engine does not calculate the KV cache for the requested context: leave some headroom.

02
Estimated GPU speed
tok/s = the model’s low, mid, or high value depending on the GPU class

Heuristic: no measurements on your machine. The Apple profiles use their unified-memory budget.

03
CPU fallback
≤3B: mid × 0.3; ≤8B: low × 0.4; otherwise max(1, low × 0.15)

Rounded to tokens/s. This branch is also used when a GPU is present but does not contain the weights.

Where do the data

Everything is sourced. Nothing is pulled out of thin air. If any information is false or outdated, signalez-la.

Hugging Face
Model cards, sizes, architectures
huggingface.co
Ollama library
Available quantizations and weights
ollama.com/library
llama.cpp
PPL benchmarks, supported architectures
github.com/ggerganov/llama.cpp
TechPowerUp
GPU specs (VRAM, bandwidth)
techpowerup.com/gpu-specs
Our test bench
Three models, one RTX 5070 Ti Laptop, published raw data
see Benchmarks

Known biases

!
What to know before reading our figures
  • The estimated tokens/s assume a machine at rest, without Chrome with 40 tabs open.
  • The September 12, 2026 driver runs with Ollama 0.30.6; its options and versions are archived.
  • The configurator does not model Macs precisely — Apple Silicon behaves nonlinearly with very large models.
  • The engine selects the highest quantization that fits within the weight budget. This choice guarantees neither loading with a long context nor quality for your use case.