GIGABYTE AI TOP ATOM: review for local AI
GIGABYTE's AI mini PC: NVIDIA GB10 chip and 128 GB of unified memory in one liter. We look at what matters for running LLMs locally — memory, bandwidth, speed, price — and who it is (or is not) the right buy for.
Specs
| Memory | 128 GB unified (LPDDR5x), 121.7 GB seen by the system |
| Bandwidth | 273 GB/s |
| Compute | Blackwell GPU and 20 Arm cores; up to 1 PFLOP in FP4 with sparsity according to NVIDIA (announced, not measured by us) |
| Power | 240 W USB-C adapter; GPU at 10.4 W at idle and 39 to 53 W on average under load in our readings (GPU only, no wattmeter) |
| Price | €6,346.27 (ATAGB10-9001, 4 TB PCIe 4.0, official AORUS store, sold out; in stock on Amazon.fr) · €7,999.95 (ATAGB10-9000, 4 TB PCIe 5.0, LDLC and Materiel.net, out of stock) — prices checked in France on October 8, 2026 · 64 GB version announced by NVIDIA from $4,999 (United States), available from October 23, 2026; GIGABYTE confirmed a 64 GB AI TOP ATOM on October 5, 2026 |
| Largest model | Qwen3-235B-A22B at 3 bits (90.5 GB used, 21 GB of margin) · Llama 3.3 70B at about 5 tokens per second, 2.6 to 4 times faster with a small draft model · gpt-oss 120B at 58 tokens per second; our 13 models, from 4B to 235B, fit with a 32,768-token context |
Who is it for?
CUDA developers, teams that want a private AI assistant, local fine-tuning, large MoE models of more than 100B.
Pros and cons
Pros:
- 128 GB of unified memory: our 13 models, from 4B to 235B (at 3 bits), fit with a 32,768-token context
- The speeds of NVIDIA's DGX Spark, within 4.3% at most with llama.cpp and within 2% with Ollama
- Stable speed under continuous load: 0.0% drift over 1 h at 16 users, −0.6% over 2 h at 8 users, no throttling reported
- 4 TB PCIe 5.0 SSD on the ATAGB10-9000: 13 GB/s read
- Complete CUDA stack: llama.cpp, vLLM, TensorRT-LLM and Unsloth all worked for us
- One-liter chassis, 1.2 kg, clean front
Cons:
- Priced above a 128 GB Ryzen AI Max+ 395 mini PC
- A dense 70B model generates at about 5 tokens per second, a limit set by the GB10 platform's memory bandwidth (273 GB/s): an MoE model or a small draft model (×2.6 to ×4) gives conversational speeds
What we measured

GIGABYTE kindly provided us with an AI TOP ATOM (ATAGB10-9000: 128 GB, 4 TB PCIe 5.0 SSD) for a series of guides. The measurements and opinions are our own; GIGABYTE did not review or approve this content before publication. The method (protocol frozen before the first measurement), the model file hashes and the data behind our tables are published on our methodology page. Our data, charts and photos are free to reuse under the CC BY 4.0 license, crediting bestllmfor.com.
| Guide | What it shows |
|---|---|
| Which LLMs really run | 13 models from 4B to 235B: all fit with a 32,768-token context; gpt-oss 120B at 58 tokens per second |
The series has 7 guides, published one per week: this page grows with each new guide.
AI TOP ATOM model numbers
GIGABYTE offers the ATOM in six model numbers, from ATAGB10-9000 to ATAGB10-9005. The first three have 128 GB: 9000 with a 4 TB PCIe 5.0 SSD (the one we measured), 9001 with 4 TB PCIe 4.0, 9002 with 1 TB PCIe 4.0. A 64 GB version is announced from October 23, 2026. Check the exact model number with the reseller: some listings misdescribe the SSD.
Who it's for
- A team of 5 to 20 people that wants a private AI assistant on its documents (measured up to 10 simultaneous users).
- A CUDA developer who wants the full NVIDIA stack on their desk.
- Local fine-tuning: LoRA of a Qwen3-8B, QLoRA of a 70B and of gpt-oss 120B with Unsloth.
- Large MoE models of more than 100 billion parameters, at comfortable speed.
Verdict
The AI TOP ATOM puts NVIDIA's GB10 platform and its 128 GB on a desk, in a one-liter, 1.2 kg case, with speeds within 4.3% of those published for the DGX Spark. It excels on large MoE models (gpt-oss 120B at 58 tokens per second), on long documents, when several people use it and for local training. For a team that wants a private AI assistant or a developer rooted in CUDA, it is an excellent choice. If you work alone and generate with small dense models, or if budget decides, also compare the Mac M5 Max and Ryzen AI Max+ 395 mini PCs.
FAQ
Can the GIGABYTE AI TOP ATOM run a 70B LLM locally?
Qwen3-235B-A22B at 3 bits (90.5 GB used, 21 GB of margin) · Llama 3.3 70B at about 5 tokens per second, 2.6 to 4 times faster with a small draft model · gpt-oss 120B at 58 tokens per second; our 13 models, from 4B to 235B, fit with a 32,768-token context.
How much does the GIGABYTE AI TOP ATOM cost?
Expect €6,346.27 (ATAGB10-9001, 4 TB PCIe 4.0, official AORUS store, sold out; in stock on Amazon.fr) · €7,999.95 (ATAGB10-9000, 4 TB PCIe 5.0, LDLC and Materiel.net, out of stock) — prices checked in France on October 8, 2026 · 64 GB version announced by NVIDIA from $4,999 (United States), available from October 23, 2026; GIGABYTE confirmed a 64 GB AI TOP ATOM on October 5, 2026. Prices move fast — check today's price with retailers.
Who is the GIGABYTE AI TOP ATOM for?
CUDA developers, teams that want a private AI assistant, local fine-tuning, large MoE models of more than 100B.