Intermediate 21 minMini-PC

GB10 128 GB: which LLMs actually run (mesures)

Direct response

On a GIGABYTE AI TOP ATOM (NVIDIA GB10, 128 GB), all 13 measured models, from 4B to 235B, fit with a 32,768-token context and at least 20 GB of headroom. Large MoE models, which activate only part of their parameters for each token, are much faster than dense models of the same size: gpt-oss 120B outputs 58 tokens per second. A 70B dense model outputs 4.8 tokens per second: memory bandwidth, not capacity, determines the speed.

One hundred twenty-eight gigabytes of unified memory: that is the advantage GB10 machines have over graphics cards. But what can you really load into it, with how much headroom and at what speed? We measured 13 models, from 4B to 235B, on an AI TOP ATOM graciously provided by GIGABYTE, using a protocol fixed before the first measurement. Here is what fits, what feels good for everyday use, and who should buy it.

By Mohamed Meguedmi·Update 2026-10-08·Measured on GIGABYTE AI TOP ATOM
i
Transparency
Hardware provided free of charge by GIGABYTE for this series of guides. The measurements and opinions are ours; GIGABYTE neither reviewed nor approved this content before publication. This page contains no affiliate links. Our data, infographics, and photos may be freely reused under the CC BY 4.0 license, with attribution to BestLLMfor.
Key figures
13models
from 4B to 235B, all loaded with 32,768 tokens of context
58tok/s
written by gpt-oss 120B, a 117-billion-parameter model
160,3tok/s
in total for 16 simultaneous users on gpt-oss 120B
20GB
minimum headroom for each of the 13 models, including context
The GIGABYTE AI TOP ATOM viewed from the front-right at a three-quarter angle, lying flat: anthracite-gray cover and black slatted front panel.
GIGABYTE’s AI TOP ATOM: 1,2 kg and about one liter for a NVIDIA GB10 chip and 128 GB of unified memory. Photo of our unit, with the background neutralized. Enlarge ↗

#The tested machine

Tested machine specifications: GIGABYTE AI TOP ATOM ATAGB10-9000, measured on 2026-10-06: NVIDIA GB10 chip (Grace Blackwell), 20 Arm cores (10 X925 and 10 A725), 128 GB of LPDDR5x memory at 273 GB/s, 4 TB PCIe 5.0 SSD (32 GT/s, 4 lanes), 10 GbE network and 2 QSFP ports (ConnectX-7, 200 Gb/s), 150 × 150 mm form factor, 1 liter and 1,200 g, 240 W USB-C power supply, DGX OS 7.5.0
The configuration collected by script on our unit before the first measurement, supplemented by the GIGABYTE specifications for the form factor and ports. The SSD's PCIe 5.0 link was verified (32 GT/s, 4 lanes). Colors: blue for the chip (GB10 and 20 Arm cores); green for memory (128 GB) and storage (4 TB SSD); gray for networking, form factor, power, and software. Enlarge ↗

The machine is a GIGABYTE AI TOP ATOM, model ATAGB10-9000: NVIDIA GB10 chip (20 Arm cores and a Blackwell GPU sharing 128 GB of memory advertised at 273 GB/s), with a 4 TB PCIe 5.0 SSD. It runs DGX OS 7.5.0, the NVIDIA system based on Ubuntu (driver 580.159.03, CUDA 13.0).

Three software tools were used. llama.cpp, a widely used open-source engine for running models locally, compiled on-site for the GB10 chip, provides the main results table. Ollama 0.34.1 is used for comparison with a MacBook Pro M5 Max measured with the same version. vLLM 26.09, a server designed to serve multiple people simultaneously, runs in the container provided by NVIDIA.

The AI TOP ATOM black box in its shipping carton, secured with foam, bearing the GIGABYTE AI TOP and Accelerated by NVIDIA markings.
Unboxing: the AI TOP ATOM arrives in its box, cushioned with foam. Our unit was received on September 29, 2026. Enlarge ↗

#How we measured

The llama.cpp speeds come from five repetitions, with a standard deviation (the spread between repetitions) of at most 3%. Load times, reading a 30,000-token document, and memory headroom are single measurements; multi-user tests are repeated three times. Before the first measurement, a three-minute test verified that the GPU could sustain its compute power.

The files for the 13 models have a SHA-256 fingerprint (a digital signature for the file) identical to the one published by Hugging Face, the site that hosts these models. The protocol, written on October 5, was locked on the 6th at 13:30, before the first measurement.

An amendment issued the same day, around 16:50, moved the cold remeasurement to the evening and added the tests reproduced here: the official Ollama script, speculative decoding, sustained load, and Llama 3.3 70B in NVFP4. The method, file hashes, and data from our tables are published on our methodology page.

The main figures were remeasured cold that same evening, after 20 minutes of idle time and with the cache cleared: at most a 2.6% difference, below the protocol's 3% threshold.

#The same speed as a DGX Spark, within a few percent

To our knowledge, 128 GB GB10 machines share the same NVIDIA chip and the same memory; their construction, cooling, and storage may differ. To position the ATOM, we reran two published benchmarks for NVIDIA’s DGX Spark, using the same software version and settings.

With llama.cpp (build 7941, the one from the table published by the project; our main table uses the newer build 11430), across six common models, read performance is within 2.3%; write performance is 2.8% lower at the median and 4.3% lower at most. With the official Ollama script (version 0.12.6, the one used for its published measurements), our three measurements are within 2%, for both reading and writing: gpt-oss 20B, gpt-oss 120B, and Llama 3.1 70B.

These guide speeds should therefore hold, within a few percent, for the other 128 GB GB10 machines; we measured only the ATOM. Its construction and sustained-load behavior (temperatures, stability over 1 hour and 2 hours) are detailed in the next guide in the series, and its power consumption in a dedicated guide.

#13 models, from 4B to 235B: footprint and speed

→
Read, write, tokens, and GB
Read speed indicates how quickly the machine processes your question (the “prompt”) or your document before responding; write speed indicates how quickly the response appears. The first matters for long documents, code, and agents, while the second affects comfort. Both are expressed in tokens per second (in French, a token is about two-thirds of a word). Memory is measured in GB, as Linux displays it: 1 GB = 1 024 MB.

The table summarizes the test run. “Memory used” is the drop in available memory once the model is loaded with 32,768 context tokens reserved. “Load time” is the cold-start loading time with the disk cache cleared. The speeds come from the llama-bench measurement tool (llama.cpp, build 11430 from October 5, 2026), first with an empty context and then with 32,768 tokens already present; actual times for a long document are lower.

13 models measured on the AI TOP ATOM, October 6, 2026 (llama.cpp b11430, DGX OS 7.5.0, driver 580.159.03). Writing and reading in tokens per second.
ModelTypeMemory usageChargementWriting (empty → 32k)Reading (empty → 32k)Usage
Gemma 3 4B (Q4_0)dense4.6 GB3 s80,8 → 63,66 239 → 5 409very smooth
Qwen2.5-Coder 7B (Q8_0)dense10.2 GB3 s30,0 → 23,33 746 → 2 138fluide
gpt-oss 20B (MXFP4)MoE, 3.6B active13.0 GB4 s81,4 → 62,54 950 → 3 316very smooth
Qwen3.8 27B (Q4_K_XL)dense19.1 GB5 s11,8 → 10,6865 → 717correct
Qwen3.6 35B-A3B (Q4_K_XL)MoE, 3B active22.4 GB5 s66,0 → 55,62 987 → 2 429very smooth
GLM-4.7-Flash (Q8_0)MoE, 3B active32.6 GB6 s51,4 → 35,52 392 → 608very smooth
Qwen3-Coder 30B-A3B (Q8_0)MoE, 3.3B active34.3 GB6 s62,6 → 33,23 377 → 1 603very smooth
Llama 3.3 70B (Q4_K_M)dense51.1 GB8 s4,8 → 3,9405 → 269ideal for batch processing
gpt-oss 120B (MXFP4)MoE, 5.1B active61,7 GB10 s58,0 → 42,22 609 → 1 832very smooth
Qwen3.5 122B-A10B (Q4_K_XL)MoE, 10B active74.5 GB12 s23,1 → 21,41 126 → 945fluide
Nemotron-3 Super 120B-A12B (Q4_K_XL)MoE, 12B active80.4 GB12 s16,9 → 16,5851 → 809correct
Qwen3.8-Flash-Next 125B (IQ4_XS)MoE, approximately 6B active89.5 GB21 s27,3 → 25,21 073 → 917fluide
Qwen3-235B-A22B (Q2_K_XL)MoE, 22B active90.5 GB12 s17,7 → 11,8588 → 331acceptable; heavy compression (Q2)

To read the table: a dense model uses all its parameters for every token, while an MoE model (“mixture of experts”) uses only a fraction, indicated in billions (B). The acronym in parentheses refers to weight compression. In our files, Q8 uses 8.5 bits per parameter, Q4 and MXFP4 formats use 4.3 to 5.6 bits, and Q2_K_XL uses 3 bits: it is the strongest compression in the table.

The “Usage” column applies our guidelines to write speed with an empty context: very smooth above 40 tokens per second, smooth from 20 to 40, acceptable from 10 to 20. Below that, we reserve the model for batch processing.

Bar chart of the writing speed of the 13 models with an empty context: from 81.4 tokens per second for gpt-oss 20B to 4.8 for Llama 3.3 70B; MoE models in orange, dense models in blue
Write speed in tokens per second, with an empty context. In orange, with the two-box icon lit: MoE models, which activate only part of their parameters for each token. In blue, with the solid icon: dense models. At comparable sizes, MoE models are much faster. Enlarge ↗

First takeaway: nothing in this table pushes the machine to its limits; even Qwen3-235B leaves 21 GB free with a 30,000-token context. Second takeaway, more useful for choosing: capacity is almost never the constraint; the model choice determines speed, from 4.8 to 81.4 tokens per second.

#What determines speed: active parameters, not size

To write each token, the chip must reread from memory the weights used for that token. With 273 GB/s of bandwidth, the maximum speed is easy to calculate: 273 divided by the amount of weights read for each token. A dense model reads all its weights every time; an MoE model reads only a fraction, the “experts” selected for that token.

That explains the table's paradox. The gpt-oss 120B file weighs 59 GB, but the model activates only 5.1 billion parameters per token, so it generates at 58.0 tokens per second. Qwen3.8 27B, whose file is more than three times lighter (16 GB), is a dense model: it activates all 27 billion parameters for every token and generates at 11.8 tokens per second. The larger model is nearly five times faster than the smaller one.

Measured write speed versus the theoretical ceiling (273 GB/s ÷ active-weight volume, both in decimal units: 1 GB = 1 billion bytes), with an empty context. Calculated by us as an order-of-magnitude estimate. Active parameters: model specifications for MoE models; number reported by llama.cpp for dense models. Shares calculated from unrounded values.
ModelActive parametersTheoretical ceilingMeasuredShare of the cap
Qwen2.5-Coder 7B (dense)7.6B33,730,089 %
Llama 3.3 70B (dense)70.6B6,44,874 %
Qwen3.8 27B (dense)27.3B15,611,876 %
Qwen3-Coder 30B-A3B (MoE)3.3B77,862,681 %
gpt-oss 120B (MoE)5.1B98,758,059 %
Qwen3.6 35B-A3B (MoE)3B141,166,047 %

The table includes six models, those whose active parameter count is published or reported by llama.cpp. Dense models reach 74 to 89% of that ceiling: the GB10 makes use of almost all its memory.

Across all thirteen models, the eight MoE models other than Qwen3.8-Flash-Next reach 47% to 81%, with six of them between 52% and 62%; selecting experts and performing auxiliary computations probably adds a fixed amount of time to each token. Qwen3.8-Flash-Next is excluded from this calculation: it has 51 billion additional lookup-table parameters on top of its 125 billion, making the estimate unsuitable.

At comparable sizes, MoE models are still much faster. Hence our main recommendation: on a GB10 machine, favor an MoE for a large model. A small dense model such as Gemma 3 4B also writes very quickly (80.8 tokens per second), but it is much smaller and serves different use cases.

#Models with more than 100 billion parameters

That's the purpose of 128 GB. NVIDIA advertises models of up to 200 billion parameters for the platform; our measurements confirm that claim, and a compressed 235B at 3 bits even fits beyond that. Five models with more than 100 billion parameters fit on the ATOM, all with at least 20 GB of headroom at a 30,000-token context.

gpt-oss 120B, the best speed-to-footprint compromise
58.0 tokens per second, using 61.7 GB with its context, loaded in 10 seconds. OpenAI publishes it directly in MXFP4 format: it fits without additional compression.
Qwen3.8-Flash-Next 125B, the fastest of the giant Qwen models
27.3 tokens per second with approximately 6 billion active parameters, occupying 89.5 GB in IQ4_XS. Its less-compressed Q4_K_XL version also fits, but with a narrow margin (see below).
Qwen3.5 122B-A10B
23.1 tokens per second, 74.5 GB; its ten billion active parameters place it behind gpt-oss.
Nemotron-3 Super 120B-A12B (NVIDIA)
16.9 tokens per second, 80.4 GB. With twelve billion active parameters, it writes more slowly than gpt-oss, and it is the model whose writing holds up best as the context grows: 16.5 at 32,768 tokens, about 3% less.
Qwen3-235B-A22B, separately
It fits in Q2_K_XL, the table's strongest compression (3 bits per parameter on average): 17.7 tokens per second, 90.5 GB. Such strong compression generally reduces response quality; we did not measure it.

#Where memory runs out: around 100 to 105 GB of weights

The system sees 121.7 GB of memory: some of the 128 GB is reserved at startup. Once the machine was running, with nothing else in memory, 108 to 113 GB remained available depending on the time. To find the ceiling, we loaded two heavier versions of the largest models, with the same 32,768-token context and a safeguard that stops loading if free memory falls below 3 GB.

The two heaviest loads, both successful, October 6, 2026 (llama-server b11430, 32,768-token context). Headroom: memory still available after reading a 30,000-token document. “Just”: 5 to 15 GB of headroom.
ModelFileMemory usageRemaining headroomWriting (after reading 4,000 → 30,000 tokens)Place
Qwen3.8-Flash-Next 125B (Q4_K_XL)103.7 GB107.0 GB6.0 GB24.9 → 21.9 tok/sjust
Qwen3-235B-A22B (Q3_K_XL, 3.5 bits)97.0 GB104.7 GB8.2 GB13.9 → 10.4 tok/sjust

No model failed to load, but these two leave little headroom: barely enough room for a second model, a much longer context, or a demanding application running alongside them. The practical limit is therefore around 100 to 105 GB of weights. This excludes, for example, Qwen3.8-Flash-Next's NVFP4 weights (a compressed format from NVIDIA): about 135 GB according to the Kubesimplify blog (August 27, 2026), which notes that two machines are then required.

That same post had already shown the model running on a single machine in GGUF (the llama.cpp format), but only the most compressed version existed at the time. On the ATOM, the IQ4_XS and Q4_K_XL versions run, with 21 and 6 GB of headroom.

→
The right size for comfortable use
Aim for no more than about 90 GB for your context: that leaves about twenty GB for the system, a second small model, or a longer context. Our thirteen models stay within that guideline.

#The cost of long context

A long document, a codebase, or a conversation that keeps going: every token already present in the context slows down what follows. With 32,768 tokens in memory, writing speed drops by 3 to 47% depending on the model. We also measured the real-time needed to read a 30,000-token document in one block, about fifty pages.

Time to read a 30,000-token document in one pass (llama-server b11430, October 6, 2026), followed by the response generation speed.
ModelReading 30,000 tokensWriting afterward
Gemma 3 4B4,7 s60.8 tok/s
gpt-oss 20B8,1 s61.6 tok/s
Qwen2.5-Coder 7B12,7 s23.3 tok/s
Qwen3.6 35B-A3B14,2 s54.4 tok/s
Qwen3-Coder 30B-A3B17,4 s33.3 tok/s
gpt-oss 120B21,8 s42.8 tok/s
GLM-4.7-Flash32,7 s35.6 tok/s
Qwen3.8 27B39,4 s10.6 tok/s
Qwen3.8-Flash-Next 125B42,9 s23.0 tok/s
Qwen3.5 122B-A10B43,6 s21.2 tok/s
Nemotron-3 Super 120B-A12B55,4 s16.2 tok/s
Qwen3-235B-A22B1 min 33 sec12.1 tok/s
Llama 3.3 70B1 min 44 sec4.0 tok/s

For most models, these times are longer than the « Read » column in the first table suggests. The likely reason: llama-server, the server used in practice, processes text in batches of 512 tokens by default, versus 2 048 in our llama-bench configuration.

Write-speed curves according to the context already present, from 0 to 32,768 tokens: one color per model: gpt-oss 120B from 58.0 to 42.2, Qwen3.6 35B-A3B from 66.0 to 55.6, Qwen3-Coder 30B-A3B from 62.6 to 33.2, Nemotron-3 Super 120B from 16.9 to 16.5, Qwen3.8 27B from 11.8 to 10.6, Llama 3.3 70B from 4.8 to 3.9
Writing speed in tokens per second based on the context already present, from 0 to 32,768 tokens (horizontal axis in thousands of tokens). One color per model, with its name written to the right of each curve; the document icon, with “→ 32,768,” indicates the maximum measured context. Nemotron (−3%) and Qwen3.8 27B (−10%) retain almost all their speed; with 32,768 tokens in memory, gpt-oss 120B still writes 42 tokens per second and Qwen3.6 35B-A3B nearly 56. Enlarge ↗

Reading is a strong point for the GB10: compared with the Mac tested below, it reads gpt-oss 120B 31% faster. If you submit a report of around fifty pages, the response starts 22 seconds later with gpt-oss 120B, and 1 min 44 sec later with a dense 70B model. This criterion carries significant weight for document analysis and coding agents.

#Up to 16 concurrent users: what the machine can handle

With vLLM, we simulated 1, 8, and then 16 concurrent users. Each sends a request for 1,024 tokens and receives a response of 512 tokens; each point is measured three times with different requests, and we publish the median.

Multiple simultaneous users, vLLM 26.09 (container NVIDIA), October 6, 2026. “Per person”: writing speed once the response has started. “Total”: tokens written per second for all users combined, including the wait for the first word. “Slowest cases”: the 99th percentile, the time under which 99 out of 100 requests finish, calculated over 4 to 64 requests per batch, so it is close to the maximum observed.
ModelUsersTotalPer personFirst token (average)First token (slowest cases)
gpt-oss 120B135.6 tok/s36.4 tok/s0,34 s0,35 s
gpt-oss 120B8113.9 tok/s14.5 tok/s0,97 s1,62 s
gpt-oss 120B16160.3 tok/s10.2 tok/s1,07 s3,71 s
gpt-oss 20B149.1 tok/s50.0 tok/s0,16 s0,16 s
gpt-oss 20B8205.9 tok/s26.5 tok/s0,49 s0,81 s
gpt-oss 20B16322.4 tok/s20.7 tok/s0,52 s1,73 s
Total and per-user throughput curves for 1, 8, and 16 concurrent users: gpt-oss 120B goes from 35.6 to 160.3 tokens per second in total, 10.2 per user at 16; gpt-oss 20B goes from 49.1 to 322.4 in total, 20.7 per user at 16
Tokens-per-second throughput for 1, 8, and 16 simultaneous users (vLLM). Purple: gpt-oss 20B; orange: gpt-oss 120B. Below the horizontal axis, a silhouette represents one user, while a group represents multiple users. Solid lines, group icon: total throughput. Dotted lines, person icon: speed seen by each person. At 16 users, gpt-oss 20B exceeds 320 tokens per second in total. Enlarge ↗

Between 1 and 16 users, total throughput is multiplied by 4.5 with gpt-oss 120B and by 6.6 with gpt-oss 20B: when several requests share the reading of the same weights, the chip fully leverages its computing power.

With 16 people, each person still sees their response being generated at 10 tokens per second with gpt-oss 120B, and at 21 with gpt-oss 20B. With the 120-billion-parameter model, the first word arrives in just over one second on average, and in nearly 4 seconds in the slowest cases.

For a single person, however, vLLM isn't the fastest: without performance tuning, it generates 36.4 tokens per second on gpt-oss 120B, compared with 54.4 for llama.cpp on long responses. Its logs show that on the GB10 it selects the Marlin compute kernel for the MXFP4 format, which probably explains the difference. vLLM makes sense as soon as multiple people or agents share the machine.

#Alone at the machine: Ollama or llama.cpp?

Ollama is the simplest way to get started, and it runs on the GB10 without configuration. On gpt-oss, however, it is not the fastest. For long responses, version 0.34.1 writes at 42.3 tokens per second on gpt-oss 120B, compared with 54.4 for llama.cpp in the same MXFP4 format, or 22% less. On gpt-oss 20B: 58.8 versus 78.8, or 25% less.

With Qwen models, it's the opposite. On Qwen3.8 27B, Ollama writes between 22,9 and 30,5 tokens per second depending on the passage, versus 11,8 for llama.cpp with its default settings; on Qwen3.6 35B-A3B, between 90,5 and 94,9, versus 66,0. The likely reason: Ollama enables speculative decoding by default, proposing several tokens ahead that the model validates all at once (the draft_num_predict setting appears in these models' configuration).

Our test points in the same direction. On llama-server (six 400-token texts, prose and code), llama.cpp's speculative decoding (MTP, which predicts several tokens at once) takes Qwen3.8 27B from 11.7 to 21.5 tokens per second in prose, and from 11.6 to 27.2 in code.

→
Ollama enables long contexts by default
Without any configuration on our part, Ollama 0.34.1 opened contexts of 131,072 tokens for gpt-oss and 262,144 for Qwen and Gemma, even though its documentation indicates 4,096 by default. The memory it reports remains close to our measurements. To leave room for a second model, reduce the context with the OLLAMA_CONTEXT_LENGTH variable.

In practice: Ollama to get started and for Qwen models, which it accelerates out of the box; llama.cpp to get the most out of gpt-oss, or Qwen3.8 with speculative decoding enabled. The best choice depends more on the settings than the machine.

#Compared with the MacBook Pro M5 Max 128 GB

A reader measured their 128 GB MacBook Pro M5 Max (40-core GPU) on September 16 with Ollama 0.34.1, using a single pass per model; their measurements are published in our dedicated guide. We reran the same series on the ATOM, using two passes: the same version of Ollama, the same models, and the same prompt.

Same version of Ollama (0.34.1), same models, same prompt. Writing in tokens per second; reading a document of about 17,000 tokens for the final line. Mac: one pass, measured by a reader on September 16, 2026. ATOM: two passes, on October 6, 2026.
MetricATOM, 1st passATOM, 2nd passMacBook Pro M5 MaxGap
Writing, gemma4:12b45,954,158,1Mac +7 to +27%
Writing, qwen3.8:27b30,522,936,6Mac +20 to +59%
Writing, gpt-oss:20b58,158,8113,4Mac +93 to +95%
Writing, gpt-oss:120b42,242,379,1Mac +87%
Reading, gpt-oss:120b1 816—1 388ATOM +31%

The Mac writes faster on all four models, nearly twice as fast on the two gpt-oss models, whose runs agree: its chip has 614 GB/s of bandwidth, more than twice that of the GB10. On gemma4 and qwen3.8, the gap varies from one run to another; speculative decoding, whose gain depends on the generated text, is one possible explanation.

The ATOM reads faster, probably thanks to its GPU’s computing power, on similar but non-identical texts (16,850 and 16,689 tokens). With llama.cpp, the generation gap narrows: 54.4 tokens per second on gpt-oss 120B, versus 79.1 on the Mac under Ollama.

#Who ATOM is the right choice for

The following is based on our measurements on the ATOM.

ATOM is the right choice if you want large models at home
Five models with more than 100 billion parameters fit comfortably. To our knowledge, no consumer graphics card comes close to this 128 GB.
… if you work with long documents or code
30,000 tokens read in 22 seconds with gpt-oss 120B.
… if several people or agents share it
160 tokens per second total for 16 users on gpt-oss 120B.
… if you want the NVIDIA ecosystem
CUDA 13, vLLM in the NVIDIA container, and llama.cpp compiled for the GB10 chip worked for us on DGX OS 7.5.0.
If you're alone and primarily want writing speed
Compare this with the MacBook Pro M5 Max: thanks to its 614 GB/s, it writes faster with gpt-oss, while the ATOM reads a long document 31% faster on gpt-oss 120B. Compare the prices too.
If your models stay under 35 GB
The 64 GB version of the ATOM, announced for October 23, 2026, should be sufficient (calculation based on our measurements, not measured on this version).

#64 or 128 GB?

On October 2, 2026, NVIDIA announced 64 GB DGX Spark systems for models of up to 100 billion parameters, available on October 23 from Acer, ASUS, Dell, GIGABYTE, HP, and MSI, starting at $4,999. GIGABYTE confirmed an AI TOP ATOM 64 GB with the same design on October 5. We have not measured it.

Our measurements allow a calculation. Models up to Qwen3-Coder 30B-A3B take up less than 35 GB with 32,768 tokens of context and would fit; Llama 3.3 70B (51 GB) would be borderline. gpt-oss 120B (62 GB) and larger models would not fit. At equal bandwidth, which still needs to be verified, speeds should be close.

For a model with more than 100 billion parameters on a single machine, the 128 GB version is the clear choice. According to NVIDIA, two 64 GB machines connected together can also pool their memory; we have not tested this.

#The prices recorded

On October 8, 2026, the 4 TB PCIe 4.0 version (ATAGB10-9001, same GB10 chip and same 128 GB) was listed at €6,346.27 including tax on the official AORUS store, sold out; it was in stock on Amazon.fr, sold by Amazon UK (one unit). The tested version, 4 TB PCIe 5.0 (ATAGB10-9000), was priced at €7,999.95 at LDLC and Materiel.net, out of stock.

Prices and stock for these machines change quickly: check them before buying.

#Our verdict

The ATOM is an excellent choice for running a 120-billion-parameter model at home, sharing it with a team, or quickly reading long documents. It delivered the platform’s reference performance, with no thermal throttling reported in our measurements, including during an hour of continuous load with 16 users. If budget matters, its PCIe 4.0 version retains the same chip and memory.

#What we did not measure

Response quality
This guide measures what fits and how fast it runs, not how good each model is.
Temperatures, noise, and power consumption
The next guide in the series covers temperatures under sustained load; power consumption has its own guide, measured at the chip (GPU only). For noise, we cite measurements from Hardware & Co.
Contexts beyond 32,768 tokens
Several models support much longer contexts; this guide stops at 32 768 tokens, and the series' 70B tutorial goes up to 131 072 tokens for Llama 3.3 70B and gpt-oss 120B.
The other GB10 machines and the 64 GB version
Our figures match those published for the NVIDIA DGX Spark, but we measured only the 128 GB ATOM.

Updates and fixes: this page will be revised if a new version of DGX OS, llama.cpp, or Ollama changes these results; each correction will be dated.

#FAQ

FAQ
What is the largest model you can run on a 128 GB GB10 machine?+
On our AI TOP ATOM, the largest model tested was Qwen3-235B-A22B: 90.5 GB used in Q2_K_XL (3 bits per parameter), with 21 GB of headroom at a 30,000-token context. Its Q3_K_XL version also fits, with 8 GB of headroom. We did not measure response quality: at 3 bits, a model deviates more from its original version than at 4 or 5 bits, but that does not indicate which one responds better.
Can a 70B model be used daily on a GB10?+
It runs without difficulty (51 GB occupied) and writes at 4.8 tokens per second; it reads 30,000 tokens in 1 min 44 sec. It is ideal for batch processing: in NVFP4 form under vLLM, eight simultaneous requests total 37.7 tokens per second, and each response, once started, is written at 5.0 tokens per second. A small draft model brings it to 12 tokens per second (see our 70B tutorial). In conversation, gpt-oss 120B is much more pleasant.
Do this guide's speeds apply to a NVIDIA DGX Spark or another brand?+
Most likely, within a few percent, but we measured only the ATOM. To our knowledge, 128 GB GB10 machines share the same chip and memory. Replaying two published benchmarks for the DGX Spark, with llama.cpp and Ollama, our read speeds are within 2.3%, and our write speeds are at most 4.3% lower.
Should you choose the 64 GB or 128 GB version?+
If your models occupy less than 35 GB with their context, such as gpt-oss 20B or Qwen3.6 35B-A3B in our measurements, the 64 GB version should be sufficient; this is a calculation, as we did not measure it. For a model with more than 100 billion parameters on a single machine, such as gpt-oss 120B (62 GB), the 128 GB version is required.
Is Ollama the Best Choice on a GB10 Machine?+
To get started, yes: it works without configuration. On gpt-oss 120B, however, llama.cpp writes faster (54.4 versus 42.3 tokens per second on long responses). On the Qwen models, Ollama comes out ahead with its default settings, probably thanks to the speculative decoding it enables; llama.cpp comes closer when you enable its own. The best choice therefore depends mainly on the settings.
In the same series
The machineThe GIGABYTE AI TOP ATOM specs
Guide 2 · forthcomingAI TOP ATOM vs DGX Spark
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.