BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-10-08

GB10 128 GB: which LLMs actually run? 13 models measured

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

One hundred twenty-eight gigabytes of unified memory is the GB10's edge over graphics cards. But what can you really load, with what margin and at what speed? We measured 13 models, from 4B to 235B, on an AI TOP ATOM kindly provided by GIGABYTE, with a protocol frozen before the first measurement.

By Mohamed Meguedmi·Last updated 2026-10-08·21 min read·Measured on GIGABYTE AI TOP ATOM

Key takeaways

On a GIGABYTE AI TOP ATOM (NVIDIA GB10, 128 GB), the 13 models we measured, from 4B to 235B, all fit with a 32,768-token context and at least 20 GB of margin. Large MoE models, which activate only part of their parameters for each token, are much faster than a dense model of the same size: gpt-oss 120B generates 58 tokens per second. A dense 70B generates 4.8 tokens per second: it is memory bandwidth, not capacity, that sets the speed.

  • 13 models from 4B to 235B, all loaded with a 32,768-token context
  • 58 tok/s generated by gpt-oss 120B, a model with 117 billion parameters
  • 160.3 tok/s in total for 16 simultaneous users on gpt-oss 120B
  • 20 GB of margin at a minimum for each of the 13 models, context included

One hundred twenty-eight gigabytes of unified memory: that is the GB10 machines' edge over graphics cards. But what can you really load, with what margin and at what speed? We measured 13 models, from 4B to 235B, on an AI TOP ATOM kindly provided by GIGABYTE, with a protocol frozen before the first measurement. Here is what fits, what is pleasant to use day to day and who it is the right buy for.

Transparency. Hardware kindly provided by GIGABYTE for this series of guides. The measurements and opinions are our own; GIGABYTE did not review or approve this content before publication. This page contains no affiliate links. Our data, charts and photos are free to reuse under the CC BY 4.0 license, crediting bestllmfor.com.

The GIGABYTE AI TOP ATOM seen from the right at a three-quarter angle, lying flat: anthracite gray cover and black slatted front.
GIGABYTE's AI TOP ATOM: 1.2 kg and about one liter for an NVIDIA GB10 chip and 128 GB of unified memory. Photo of our unit, background neutralized.

The machine we tested

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Spec sheet of the tested machine: GIGABYTE AI TOP ATOM ATAGB10-9000, recorded on October 6, 2026: NVIDIA GB10 chip (Grace Blackwell), 20 Arm cores (10 X925 and 10 A725), 128 GB of LPDDR5x memory at 273 GB/s, 4 TB PCIe 5.0 SSD (32 GT/s, 4 lanes), 10 GbE networking and 2 QSFP ports (ConnectX-7, 200 Gb/s), 150 × 150 mm form factor, 1 liter and 1,200 g, 240 W USB-C power supply, DGX OS 7.5.0
The configuration read out by script on our unit before the first measurement, completed with GIGABYTE's spec sheet for the form factor and the ports. The SSD's PCIe 5.0 link was verified (32 GT/s, 4 lanes). Colors: blue for the chip (GB10 and 20 Arm cores); green for memory (128 GB) and storage (4 TB SSD); gray for networking, form factor, power supply and software.

The machine is a GIGABYTE AI TOP ATOM, model ATAGB10-9000: NVIDIA GB10 chip (20 Arm cores and a Blackwell GPU sharing 128 GB of memory advertised at 273 GB/s), 4 TB PCIe 5.0 SSD. It runs DGX OS 7.5.0, NVIDIA's Ubuntu-based system (driver 580.159.03, CUDA 13.0).

Three pieces of software were used. llama.cpp, a widely used open-source engine for running models locally, compiled on site for the GB10 chip, provides the main table. Ollama 0.34.1 is used for the comparison with a MacBook Pro M5 Max measured with the same version. vLLM 26.09, a server designed to answer several people at once, runs in the container provided by NVIDIA.

The AI TOP ATOM's black box in its shipping carton, cushioned by foam, with the GIGABYTE AI TOP and Accelerated by NVIDIA markings.
Unboxing: the AI TOP ATOM arrives in its box, cushioned by foam. Our unit was received on September 29, 2026.

How we measured

The llama.cpp speeds come from five repetitions, with a standard deviation (the spread between repetitions) of 3% at most. Load times, prompt processing of a 30,000-token document and memory margins are single measurements; the multi-user tests are repeated three times. Before the first measurement, a three-minute test checked that the GPU held its compute power.

The files of the 13 models have a SHA-256 hash (a digital signature of the file) identical to the one published by Hugging Face, the site that hosts these models. The protocol, written on October 5, was frozen on the 6th at 1:30 p.m., before the first measurement.

An addendum dated the same day, at around 4:50 p.m., moved the cold re-measurement up to that evening and added tests that are included here: Ollama's official script, speculative decoding, continuous load and Llama 3.3 70B in NVFP4. The method, the file hashes and the data behind our tables are published on our methodology page.

The main figures were re-measured cold the same evening, after 20 minutes of rest and with the cache cleared: a difference of 2.6% at most, below the protocol's 3% threshold.

Same speed as a DGX Spark, within a few percent

To our knowledge, 128 GB GB10 machines share the same NVIDIA chip and the same memory; their build, cooling and storage may differ. To place the ATOM, we replayed two published references for NVIDIA's DGX Spark, with the same software version and the same settings.

With llama.cpp (build 7941, the one in the table published by the project; our main table uses build 11430, which is more recent), across six common models, prompt processing is identical to within 2.3%; generation is 2.8% lower at the median and 4.3% lower at most. With Ollama's official script (version 0.12.6, the one behind its published measurements), our three measurements are within 2%, for both prompt processing and generation: gpt-oss 20B, gpt-oss 120B and Llama 3.1 70B.

The speeds in this guide should therefore hold, within a few percent, for other 128 GB GB10 machines; we only measured the ATOM. Its build and how it holds up under load (temperatures, stability over 1 h and 2 h) are detailed in the next guide in the series, and its power consumption in a dedicated guide.

13 models, from 4B to 235B: memory used and speed

Prompt processing, generation, tokens and GB. Prompt processing speed tells you how fast the machine takes in your question (the “prompt”) or your document before answering; generation speed is how fast the answer appears. The first matters for long documents, code and agents, the second for comfort. Both are expressed in tokens per second (in French, a token is worth about two thirds of a word). Memory is in GB, and 1 GB = 1,024 MB, as Linux reports it.

The table summarizes the campaign. “Memory used” is the drop in available memory once the model is loaded with 32,768 tokens of context reserved. “Load time” is the cold load time, with the disk cache cleared. Speeds come from the llama-bench measurement tool (llama.cpp, build 11430 of October 5, 2026), at empty context and then with 32,768 tokens already present; real durations for a long document are further below.

13 models measured on the AI TOP ATOM, October 6, 2026 (llama.cpp b11430, DGX OS 7.5.0, driver 580.159.03). Generation and prompt processing in tokens per second.
ModelTypeMemory usedLoad timeGeneration (empty → 32k)Prompt processing (empty → 32k)Usability
Gemma 3 4B (Q4_0)dense4.6 GB3 s80.8 → 63.66,239 → 5,409very smooth
Qwen2.5-Coder 7B (Q8_0)dense10.2 GB3 s30.0 → 23.33,746 → 2,138smooth
gpt-oss 20B (MXFP4)MoE, 3.6B active13.0 GB4 s81.4 → 62.54,950 → 3,316very smooth
Qwen3.8 27B (Q4_K_XL)dense19.1 GB5 s11.8 → 10.6865 → 717decent
Qwen3.6 35B-A3B (Q4_K_XL)MoE, 3B active22.4 GB5 s66.0 → 55.62,987 → 2,429very smooth
GLM-4.7-Flash (Q8_0)MoE, 3B active32.6 GB6 s51.4 → 35.52,392 → 608very smooth
Qwen3-Coder 30B-A3B (Q8_0)MoE, 3.3B active34.3 GB6 s62.6 → 33.23,377 → 1,603very smooth
Llama 3.3 70B (Q4_K_M)dense51.1 GB8 s4.8 → 3.9405 → 269ideal for batch processing
gpt-oss 120B (MXFP4)MoE, 5.1B active61.7 GB10 s58.0 → 42.22,609 → 1,832very smooth
Qwen3.5 122B-A10B (Q4_K_XL)MoE, 10B active74.5 GB12 s23.1 → 21.41,126 → 945smooth
Nemotron-3 Super 120B-A12B (Q4_K_XL)MoE, 12B active80.4 GB12 s16.9 → 16.5851 → 809decent
Qwen3.8-Flash-Next 125B (IQ4_XS)MoE, about 6B active89.5 GB21 s27.3 → 25.21,073 → 917smooth
Qwen3-235B-A22B (Q2_K_XL)MoE, 22B active90.5 GB12 s17.7 → 11.8588 → 331decent; heavy compression (Q2)

How to read the table: a dense model works all of its parameters on every token, while an MoE (“mixture of experts”) model works only a fraction of them, shown in billions (B). The code in parentheses refers to the compression of the weights. In our files, Q8 takes up 8.5 bits per parameter, the Q4 and MXFP4 formats 4.3 to 5.6 bits, and Q2_K_XL 3 bits: this is the strongest compression in the table.

The “Usability” column applies our benchmarks to generation speed at empty context: very smooth above 40 tokens per second, smooth from 20 to 40, decent from 10 to 20. Below that, we reserve the model for batch processing.

Bar chart of the generation speed of the 13 models at empty context: from 81.4 tokens per second for gpt-oss 20B to 4.8 for Llama 3.3 70B; MoE models in orange, dense models in blue
Generation speed in tokens per second, empty context. In orange, with the icon showing two lit squares: MoE models, which activate only some of their parameters for each token. In blue, with the solid icon: dense models. At comparable size, MoE models are much faster.

First takeaway: nothing in this table puts the machine in difficulty; even Qwen3-235B leaves 21 GB free with 30,000 tokens of context. Second takeaway, more useful when choosing: memory is almost never the limit; it is the choice of model that sets the speed, from 4.8 to 81.4 tokens per second.

What sets the speed: active parameters, not size

To generate each token, the chip must reread from memory the weights used for that token. With 273 GB/s of bandwidth, the maximum speed is simple to compute: 273 divided by the volume of weights read per token. A dense model reads all of its weights every time; an MoE model reads only a fraction, the “experts” chosen for that token.

This explains the table's paradox. The gpt-oss 120B file weighs 59 GB, but the model activates only 5.1 billion parameters per token: it generates at 58.0 tokens per second. Qwen3.8 27B, whose file is more than three times lighter (16 GB), is a dense model: it activates its 27 billion parameters on every token and generates at 11.8 tokens per second. The big model is nearly five times faster than the small one.

Measured generation speed against the theoretical ceiling (273 GB/s ÷ volume of active weights, both in decimal units: 1 GB = 1 billion bytes), empty context. Calculated by us, as an order of magnitude. Active parameters: model cards for the MoE models; count reported by llama.cpp for dense models. Shares calculated on unrounded values.
ModelActive parametersTheoretical ceilingMeasuredShare of ceiling
Qwen2.5-Coder 7B (dense)7.6B33.730.089%
Llama 3.3 70B (dense)70.6B6.44.874%
Qwen3.8 27B (dense)27.3B15.611.876%
Qwen3-Coder 30B-A3B (MoE)3.3B77.862.681%
gpt-oss 120B (MoE)5.1B98.758.059%
Qwen3.6 35B-A3B (MoE)3B141.166.047%

The table covers six models, those whose number of active parameters is published or reported by llama.cpp. Dense models reach 74 to 89% of this ceiling: the GB10 gets almost everything out of its memory.

Across all thirteen models, the eight MoE models other than Qwen3.8-Flash-Next reach 47 to 81% of it, six of them between 52 and 62%; expert selection and side computations probably add a fixed time to each token. Qwen3.8-Flash-Next is left out of this calculation: it has 51 billion lookup-table parameters on top of its 125 billion, which makes the estimate unsuitable.

At comparable size, MoE models are still much faster. Hence our main advice: on a GB10 machine, for a large model, favor an MoE. A small dense model like Gemma 3 4B also generates very fast (80.8 tokens per second), but it is much smaller and serves other uses.

Related reading

Models of more than 100 billion parameters

This is what the 128 GB is for. NVIDIA advertises models of up to 200 billion parameters for the platform; our measurements confirm that promise, and a 235B compressed to 3 bits even fits beyond it. Five models of more than 100 billion parameters fit on the ATOM, all with at least 20 GB of margin at 30,000 tokens of context.

  • gpt-oss 120B, the best compromise between speed and memory: 58.0 tokens per second, 61.7 GB used with its context, loaded in 10 seconds. OpenAI publishes it directly in MXFP4 format: it fits without additional compression.
  • Qwen3.8-Flash-Next 125B, the fastest of the giant Qwens: 27.3 tokens per second with about 6 billion active parameters, 89.5 GB used in IQ4_XS. Its less compressed Q4_K_XL version fits too, with a narrow margin (see below).
  • Qwen3.5 122B-A10B: 23.1 tokens per second, 74.5 GB; its ten billion active parameters put it behind gpt-oss.
  • Nemotron-3 Super 120B-A12B (NVIDIA): 16.9 tokens per second, 80.4 GB. With twelve billion active parameters, it generates more slowly than gpt-oss, and it is the model whose generation holds up best as the context grows: 16.5 at 32,768 tokens, about 3% lower.
  • Qwen3-235B-A22B, in a class of its own: It fits in Q2_K_XL, the strongest compression in the table (3 bits per parameter on average): 17.7 tokens per second, 90.5 GB. Compression this strong generally reduces the quality of the answers; we did not measure it.

Where memory runs out: around 100 to 105 GB of weights

The system sees 121.7 GB of memory: part of the 128 is reserved at boot. Once the machine was up, with nothing else in memory, 108 to 113 GB remained available depending on the moment. To find the ceiling, we loaded two heavier versions of the largest models, with the same 32,768-token context and a safeguard that cuts the load if free memory falls below 3 GB.

The two heaviest loads, both successful, October 6, 2026 (llama-server b11430, 32,768-token context). Margin: memory still available after processing a 30,000-token document. “Tight”: margin of 5 to 15 GB.
ModelFileMemory usedRemaining marginGeneration (after 4,000 → 30,000 prompt tokens processed)Fit
Qwen3.8-Flash-Next 125B (Q4_K_XL)103.7 GB107.0 GB6.0 GB24.9 → 21.9 tok/stight
Qwen3-235B-A22B (Q3_K_XL, 3.5 bits)97.0 GB104.7 GB8.2 GB13.9 → 10.4 tok/stight

No model failed to load, but these two leave a narrow margin: hardly any room for a second model, a much longer context or a demanding application alongside. The practical limit is therefore around 100 to 105 GB of weights. It rules out, for example, the NVFP4 weights (a compressed NVIDIA format) of Qwen3.8-Flash-Next: about 135 GB according to the Kubesimplify blog (August 27, 2026), which states that two machines are then needed.

That same post already showed the model on a single machine in GGUF (llama.cpp's format), but only the most compressed version existed at the time. On the ATOM, the IQ4_XS and Q4_K_XL versions run, with 21 and 6 GB of margin.

The right size for comfortable use. Aim for about 90 GB at most including your context: that leaves around twenty GB for the system, a second small model or a longer context. Our thirteen models respect this guideline.

The cost of long context

A long document, a codebase or a conversation that keeps stretching: every token already in the context slows what follows. With 32,768 tokens in memory, generation speed drops by 3 to 47% depending on the model. We also measured the real time to process a 30,000-token document in one go, about fifty pages.

Time to process a 30,000-token document in one go (llama-server b11430, October 6, 2026), then generation speed of the answer.
ModelProcessing 30,000 tokensGeneration afterward
Gemma 3 4B4.7 s60.8 tok/s
gpt-oss 20B8.1 s61.6 tok/s
Qwen2.5-Coder 7B12.7 s23.3 tok/s
Qwen3.6 35B-A3B14.2 s54.4 tok/s
Qwen3-Coder 30B-A3B17.4 s33.3 tok/s
gpt-oss 120B21.8 s42.8 tok/s
GLM-4.7-Flash32.7 s35.6 tok/s
Qwen3.8 27B39.4 s10.6 tok/s
Qwen3.8-Flash-Next 125B42.9 s23.0 tok/s
Qwen3.5 122B-A10B43.6 s21.2 tok/s
Nemotron-3 Super 120B-A12B55.4 s16.2 tok/s
Qwen3-235B-A22B1 min 33 s12.1 tok/s
Llama 3.3 70B1 min 44 s4.0 tok/s

For most models, these times are longer than the “Prompt processing” column of the first table suggests. The probable reason: llama-server, the server used in practice, processes text in batches of 512 tokens by default, versus 2,048 in our llama-bench configuration.

Curves of generation speed against the context already present, from 0 to 32,768 tokens: one color per model: gpt-oss 120B from 58.0 to 42.2, Qwen3.6 35B-A3B from 66.0 to 55.6, Qwen3-Coder 30B-A3B from 62.6 to 33.2, Nemotron-3 Super 120B from 16.9 to 16.5, Qwen3.8 27B from 11.8 to 10.6, Llama 3.3 70B from 4.8 to 3.9
Generation speed in tokens per second against the context already present, from 0 to 32,768 tokens (horizontal axis in thousands of tokens). One color per model, with its name written to the right of each curve; the document icon, with “→ 32,768”, marks the maximum context measured. Nemotron (−3%) and Qwen3.8 27B (−10%) keep almost all of their speed; with 32,768 tokens in memory, gpt-oss 120B still generates 42 tokens per second and Qwen3.6 35B-A3B nearly 56.

Prompt processing is a strong point of the GB10: against the Mac compared further down, it processes gpt-oss 120B 31% faster. If you drop in a report of around fifty pages, the answer starts 22 seconds later with gpt-oss 120B, and 1 min 44 s later with a dense 70B. For document analysis and coding agents, this criterion weighs heavily.

Up to 16 users at once: what the machine can handle

With vLLM, we simulated 1, 8, then 16 simultaneous users. Each sends a request of 1,024 tokens and receives a response of 512 tokens; each point is measured three times with different requests, and we publish the median.

Several simultaneous users, vLLM 26.09 (NVIDIA container), October 6, 2026. “Per person”: generation speed once the response has started. “Total”: tokens generated per second across all users, including the wait for the first word. “Slowest cases”: 99th percentile, the duration under which 99 requests out of 100 finish, computed on 4 to 64 requests per series, hence close to the maximum observed.
ModelUsersTotalPer personFirst word (average)First word (slowest cases)
gpt-oss 120B135.6 tok/s36.4 tok/s0.34 s0.35 s
gpt-oss 120B8113.9 tok/s14.5 tok/s0.97 s1.62 s
gpt-oss 120B16160.3 tok/s10.2 tok/s1.07 s3.71 s
gpt-oss 20B149.1 tok/s50.0 tok/s0.16 s0.16 s
gpt-oss 20B8205.9 tok/s26.5 tok/s0.49 s0.81 s
gpt-oss 20B16322.4 tok/s20.7 tok/s0.52 s1.73 s
Curves of total throughput and per-person throughput for 1, 8 and 16 simultaneous users: gpt-oss 120B goes from 35.6 to 160.3 tokens per second in total, 10.2 per person at 16; gpt-oss 20B goes from 49.1 to 322.4 in total, 20.7 per person at 16
Throughput in tokens per second for 1, 8 and 16 simultaneous users (vLLM). Purple: gpt-oss 20B; orange: gpt-oss 120B. Below the horizontal axis, one silhouette represents one user, a group of silhouettes several users. Solid lines, group icon: total throughput. Dotted lines, person icon: speed seen by each person. At 16 users, gpt-oss 20B exceeds 320 tokens per second in total.

From 1 to 16 users, total throughput is multiplied by 4.5 with gpt-oss 120B and by 6.6 with gpt-oss 20B: when several requests share the same reads of the model weights, the chip makes full use of its compute power.

At 16 people, each still sees the response written at 10 tokens per second with gpt-oss 120B, and at 21 with gpt-oss 20B. With the 120-billion model, the first word arrives in a little over a second on average, and in nearly 4 seconds in the slowest cases.

For a single person, however, vLLM is not the fastest: with no particular performance tuning, it generates at 36.4 tokens per second on gpt-oss 120B, against 54.4 for llama.cpp on long responses. Its logs show that on the GB10 it picks the Marlin compute kernel for the MXFP4 format, which probably explains the gap. vLLM makes sense as soon as several people or several agents share the machine.

Alone at the machine: Ollama or llama.cpp?

Ollama is the simplest way to get started, and it works on the GB10 with no tuning. On gpt-oss, though, it is not the fastest. On long responses, version 0.34.1 generates at 42.3 tokens per second on gpt-oss 120B, against 54.4 for llama.cpp in the same MXFP4 format, which is 22% less. On gpt-oss 20B: 58.8 against 78.8, which is 25% less.

On Qwen models, it is the opposite. On Qwen3.8 27B, Ollama generates between 22.9 and 30.5 tokens per second depending on the run, against 11.8 for llama.cpp with default settings; on Qwen3.6 35B-A3B, between 90.5 and 94.9, against 66.0. The probable reason: Ollama enables speculative decoding by default there, which proposes several tokens ahead that the model validates in one go (the draft_num_predict setting appears in the configuration of these models).

Our test points the same way. On llama-server (six texts of 400 tokens, prose and code), llama.cpp's speculative decoding (MTP, which predicts several tokens at once) takes Qwen3.8 27B from 11.7 to 21.5 tokens per second on prose, and from 11.6 to 27.2 on code.

Ollama opens long contexts by default. With no adjustment from us, Ollama 0.34.1 opened contexts of 131,072 tokens for gpt-oss and 262,144 for Qwen and Gemma, even though its documentation says 4,096 by default. The memory it reports stays close to our measurements. To keep room for a second model, reduce the context with the OLLAMA_CONTEXT_LENGTH variable.

In practice: Ollama to get started and for Qwen models, which it speeds up out of the box; llama.cpp to get the most out of gpt-oss, or Qwen3.8 with its speculative decoding turned on. The best choice depends on settings more than on the machine.

Versus the MacBook Pro M5 Max 128 GB

A reader measured their MacBook Pro M5 Max 128 GB (40-core GPU) on September 16 with Ollama 0.34.1, in a single run per model; the readings are published in our dedicated guide. We replayed the same series on the ATOM, in two runs: same Ollama version, same models, same prompt.

Same Ollama version (0.34.1), same models, same prompt. Generation in tokens per second; prompt processing of a document of about 17,000 tokens for the last row. Mac: one run, measured by a reader on September 16, 2026. ATOM: two runs, on October 6, 2026.
MeasurementATOM, 1st runATOM, 2nd runMacBook Pro M5 MaxDifference
Generation, gemma4:12b45.954.158.1Mac +7 to +27%
Generation, qwen3.8:27b30.522.936.6Mac +20 to +59%
Generation, gpt-oss:20b58.158.8113.4Mac +93 to +95%
Generation, gpt-oss:120b42.242.379.1Mac +87%
Prompt processing, gpt-oss:120b1,816—1,388ATOM +31%

The Mac generates faster on all four models, nearly twice as fast on the two gpt-oss models, whose runs agree: its chip has 614 GB/s of memory bandwidth, more than double the GB10's. On gemma4 and qwen3.8, the gap varies from one run to the next; speculative decoding, whose gain depends on the generated text, is one possible explanation.

The ATOM processes prompts faster, probably thanks to the compute power of its GPU, on texts that are close but not identical (16,850 and 16,689 tokens). With llama.cpp, the generation gap narrows: 54.4 tokens per second on gpt-oss 120B, against 79.1 for the Mac under Ollama.

Related reading

Who the ATOM is the right choice for

What follows comes from our measurements on the ATOM.

  • The ATOM is the right choice if you want large models at home: Five models of more than 100 billion parameters fit with room to spare. To our knowledge, no consumer graphics card comes close to these 128 GB.
  • … if you work on long documents or code: 30,000 tokens processed in 22 seconds with gpt-oss 120B.
  • … if several people or agents share it: 160 tokens per second in total for 16 users on gpt-oss 120B.
  • … if you want the NVIDIA ecosystem: CUDA 13, vLLM in NVIDIA's container and llama.cpp compiled for the GB10 chip all worked for us under DGX OS 7.5.0.
  • If you work alone and mainly care about generation speed: Compare with the MacBook Pro M5 Max: thanks to its 614 GB/s, it generates faster on gpt-oss, while the ATOM processes a long document 31% faster on gpt-oss 120B. Compare prices too.
  • If your models stay under 35 GB: The ATOM's 64 GB version, announced for October 23, 2026, should be enough (a calculation drawn from our measurements, not measured on that version).

Related reading

64 or 128 GB?

On October 2, 2026, NVIDIA announced 64 GB DGX Spark machines for models of up to 100 billion parameters, available on October 23 from Acer, ASUS, Dell, GIGABYTE, HP and MSI, starting at $4,999. GIGABYTE confirmed a 64 GB AI TOP ATOM with the same design on October 5. We did not measure it.

Our readings allow a calculation. Models up to Qwen3-Coder 30B-A3B use less than 35 GB with 32,768 tokens of context and would fit; Llama 3.3 70B (51 GB) would be at the limit. gpt-oss 120B (62 GB) and larger models would not fit. At equal bandwidth, which remains to be verified, speeds should be close.

For a model of more than 100 billion parameters on a single machine, the 128 GB version is the one to get. According to NVIDIA, two linked 64 GB machines also pool their memory; we did not test that.

Prices checked

On October 8, 2026, in France, the 4 TB PCIe 4.0 version (ATAGB10-9001, same GB10 chip and same 128 GB) was listed at €6,346.27 including VAT on the official AORUS store, sold out; it was in stock on Amazon.fr, sold by Amazon UK (one unit). The tested version, 4 TB PCIe 5.0 (ATAGB10-9000), was €7,999.95 at LDLC and at Materiel.net, out of stock.

Prices and stock for these machines change quickly: check them before you buy.

Our verdict

The ATOM is an excellent choice for running a 120-billion-parameter model at home, sharing it with a team or processing long documents quickly. It gave us the platform's reference performance, with no thermal throttling flagged in our readings, including during one hour of continuous load at 16 users. If budget matters, its PCIe 4.0 version keeps the same chip and the same memory.

Related reading

What we did not measure

  • Answer quality: This guide measures what fits and how fast, not how good each model is.
  • Temperatures, noise and power consumption: The next guide in the series details temperatures under long load; power consumption has its own guide, read from the chip (GPU only). For noise, we cite Hardware & Co's measurements.
  • Contexts beyond 32,768 tokens: Several models accept much longer contexts; this guide stops at 32,768 tokens, and the series' 70B tutorial goes up to 131,072 tokens for Llama 3.3 70B and gpt-oss 120B.
  • Other GB10 machines and the 64 GB version: Our figures line up with those published for NVIDIA's DGX Spark, but we only measured the 128 GB ATOM.

Updates and corrections: this page will be revised if a new version of DGX OS, llama.cpp or Ollama changes these results; each correction will be dated.

Related reading

Sources

FAQ

What is the largest model you can run on a 128 GB GB10 machine?

On our AI TOP ATOM, the largest one tested is Qwen3-235B-A22B: 90.5 GB used in Q2_K_XL (3 bits per parameter) and 21 GB of margin at 30,000 tokens of context. Its Q3_K_XL version fits too, with 8 GB of margin. We did not measure the quality of the answers: at 3 bits, a model strays further from its original version than at 4 or 5 bits, but that does not tell you which one answers best.

Is a 70B model usable day to day on a GB10?

It fits without difficulty (51 GB used) and generates 4.8 tokens per second; it processes 30,000 tokens in 1 min 44 s. It is ideal for batch processing: in NVFP4 under vLLM, eight simultaneous requests total 37.7 tokens per second, and each response, once started, is generated at 5.0 tokens per second. A small draft model brings it to 12 tokens per second (see our 70B tutorial). For conversation, gpt-oss 120B is much more pleasant.

Do this guide's speeds apply to an NVIDIA DGX Spark or another brand?

Very probably, within a few percent, but we only measured the ATOM. To our knowledge, 128 GB GB10 machines share the same chip and the same memory. Replaying two published references for the DGX Spark, with llama.cpp and with Ollama, our prompt processing speeds are within 2.3% and our generation speeds are at most 4.3% lower.

Should I choose the 64 GB or the 128 GB version?

If your models use less than 35 GB with their context, like gpt-oss 20B or Qwen3.6 35B-A3B in our measurements, the 64 GB version should be enough; that is a calculation, we did not measure it. For a model of more than 100 billion parameters on a single machine, like gpt-oss 120B (62 GB), the 128 GB version is required.

Is Ollama the best choice on a GB10 machine?

To get started, yes: it works with no tuning. On gpt-oss 120B, though, llama.cpp generates faster (54.4 versus 42.3 tokens per second on long responses). On Qwen models, Ollama wins with its default settings, probably thanks to the speculative decoding it enables; llama.cpp gets close when you turn on its own. The best choice therefore depends mostly on settings.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.