GB10 128 GB: which LLMs actually run? 13 models measured
One hundred twenty-eight gigabytes of unified memory is the GB10's edge over graphics cards. But what can you really load, with what margin and at what speed? We measured 13 models, from 4B to 235B, on an AI TOP ATOM kindly provided by GIGABYTE, with a protocol frozen before the first measurement.
Key takeaways
On a GIGABYTE AI TOP ATOM (NVIDIA GB10, 128 GB), the 13 models we measured, from 4B to 235B, all fit with a 32,768-token context and at least 20 GB of margin. Large MoE models, which activate only part of their parameters for each token, are much faster than a dense model of the same size: gpt-oss 120B generates 58 tokens per second. A dense 70B generates 4.8 tokens per second: it is memory bandwidth, not capacity, that sets the speed.
- 13 models from 4B to 235B, all loaded with a 32,768-token context
- 58 tok/s generated by gpt-oss 120B, a model with 117 billion parameters
- 160.3 tok/s in total for 16 simultaneous users on gpt-oss 120B
- 20 GB of margin at a minimum for each of the 13 models, context included
One hundred twenty-eight gigabytes of unified memory: that is the GB10 machines' edge over graphics cards. But what can you really load, with what margin and at what speed? We measured 13 models, from 4B to 235B, on an AI TOP ATOM kindly provided by GIGABYTE, with a protocol frozen before the first measurement. Here is what fits, what is pleasant to use day to day and who it is the right buy for.
Transparency. Hardware kindly provided by GIGABYTE for this series of guides. The measurements and opinions are our own; GIGABYTE did not review or approve this content before publication. This page contains no affiliate links. Our data, charts and photos are free to reuse under the CC BY 4.0 license, crediting bestllmfor.com.

The machine we tested
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The machine is a GIGABYTE AI TOP ATOM, model ATAGB10-9000: NVIDIA GB10 chip (20 Arm cores and a Blackwell GPU sharing 128 GB of memory advertised at 273 GB/s), 4 TB PCIe 5.0 SSD. It runs DGX OS 7.5.0, NVIDIA's Ubuntu-based system (driver 580.159.03, CUDA 13.0).
Three pieces of software were used. llama.cpp, a widely used open-source engine for running models locally, compiled on site for the GB10 chip, provides the main table. Ollama 0.34.1 is used for the comparison with a MacBook Pro M5 Max measured with the same version. vLLM 26.09, a server designed to answer several people at once, runs in the container provided by NVIDIA.

How we measured
The llama.cpp speeds come from five repetitions, with a standard deviation (the spread between repetitions) of 3% at most. Load times, prompt processing of a 30,000-token document and memory margins are single measurements; the multi-user tests are repeated three times. Before the first measurement, a three-minute test checked that the GPU held its compute power.
The files of the 13 models have a SHA-256 hash (a digital signature of the file) identical to the one published by Hugging Face, the site that hosts these models. The protocol, written on October 5, was frozen on the 6th at 1:30 p.m., before the first measurement.
An addendum dated the same day, at around 4:50 p.m., moved the cold re-measurement up to that evening and added tests that are included here: Ollama's official script, speculative decoding, continuous load and Llama 3.3 70B in NVFP4. The method, the file hashes and the data behind our tables are published on our methodology page.
The main figures were re-measured cold the same evening, after 20 minutes of rest and with the cache cleared: a difference of 2.6% at most, below the protocol's 3% threshold.
Same speed as a DGX Spark, within a few percent
To our knowledge, 128 GB GB10 machines share the same NVIDIA chip and the same memory; their build, cooling and storage may differ. To place the ATOM, we replayed two published references for NVIDIA's DGX Spark, with the same software version and the same settings.
With llama.cpp (build 7941, the one in the table published by the project; our main table uses build 11430, which is more recent), across six common models, prompt processing is identical to within 2.3%; generation is 2.8% lower at the median and 4.3% lower at most. With Ollama's official script (version 0.12.6, the one behind its published measurements), our three measurements are within 2%, for both prompt processing and generation: gpt-oss 20B, gpt-oss 120B and Llama 3.1 70B.
The speeds in this guide should therefore hold, within a few percent, for other 128 GB GB10 machines; we only measured the ATOM. Its build and how it holds up under load (temperatures, stability over 1 h and 2 h) are detailed in the next guide in the series, and its power consumption in a dedicated guide.
13 models, from 4B to 235B: memory used and speed
Prompt processing, generation, tokens and GB. Prompt processing speed tells you how fast the machine takes in your question (the “prompt”) or your document before answering; generation speed is how fast the answer appears. The first matters for long documents, code and agents, the second for comfort. Both are expressed in tokens per second (in French, a token is worth about two thirds of a word). Memory is in GB, and 1 GB = 1,024 MB, as Linux reports it.
The table summarizes the campaign. “Memory used” is the drop in available memory once the model is loaded with 32,768 tokens of context reserved. “Load time” is the cold load time, with the disk cache cleared. Speeds come from the llama-bench measurement tool (llama.cpp, build 11430 of October 5, 2026), at empty context and then with 32,768 tokens already present; real durations for a long document are further below.
| Model | Type | Memory used | Load time | Generation (empty → 32k) | Prompt processing (empty → 32k) | Usability |
|---|---|---|---|---|---|---|
| Gemma 3 4B (Q4_0) | dense | 4.6 GB | 3 s | 80.8 → 63.6 | 6,239 → 5,409 | very smooth |
| Qwen2.5-Coder 7B (Q8_0) | dense | 10.2 GB | 3 s | 30.0 → 23.3 | 3,746 → 2,138 | smooth |
| gpt-oss 20B (MXFP4) | MoE, 3.6B active | 13.0 GB | 4 s | 81.4 → 62.5 | 4,950 → 3,316 | very smooth |
| Qwen3.8 27B (Q4_K_XL) | dense | 19.1 GB | 5 s | 11.8 → 10.6 | 865 → 717 | decent |
| Qwen3.6 35B-A3B (Q4_K_XL) | MoE, 3B active | 22.4 GB | 5 s | 66.0 → 55.6 | 2,987 → 2,429 | very smooth |
| GLM-4.7-Flash (Q8_0) | MoE, 3B active | 32.6 GB | 6 s | 51.4 → 35.5 | 2,392 → 608 | very smooth |
| Qwen3-Coder 30B-A3B (Q8_0) | MoE, 3.3B active | 34.3 GB | 6 s | 62.6 → 33.2 | 3,377 → 1,603 | very smooth |
| Llama 3.3 70B (Q4_K_M) | dense | 51.1 GB | 8 s | 4.8 → 3.9 | 405 → 269 | ideal for batch processing |
| gpt-oss 120B (MXFP4) | MoE, 5.1B active | 61.7 GB | 10 s | 58.0 → 42.2 | 2,609 → 1,832 | very smooth |
| Qwen3.5 122B-A10B (Q4_K_XL) | MoE, 10B active | 74.5 GB | 12 s | 23.1 → 21.4 | 1,126 → 945 | smooth |
| Nemotron-3 Super 120B-A12B (Q4_K_XL) | MoE, 12B active | 80.4 GB | 12 s | 16.9 → 16.5 | 851 → 809 | decent |
| Qwen3.8-Flash-Next 125B (IQ4_XS) | MoE, about 6B active | 89.5 GB | 21 s | 27.3 → 25.2 | 1,073 → 917 | smooth |
| Qwen3-235B-A22B (Q2_K_XL) | MoE, 22B active | 90.5 GB | 12 s | 17.7 → 11.8 | 588 → 331 | decent; heavy compression (Q2) |
How to read the table: a dense model works all of its parameters on every token, while an MoE (“mixture of experts”) model works only a fraction of them, shown in billions (B). The code in parentheses refers to the compression of the weights. In our files, Q8 takes up 8.5 bits per parameter, the Q4 and MXFP4 formats 4.3 to 5.6 bits, and Q2_K_XL 3 bits: this is the strongest compression in the table.
The “Usability” column applies our benchmarks to generation speed at empty context: very smooth above 40 tokens per second, smooth from 20 to 40, decent from 10 to 20. Below that, we reserve the model for batch processing.
First takeaway: nothing in this table puts the machine in difficulty; even Qwen3-235B leaves 21 GB free with 30,000 tokens of context. Second takeaway, more useful when choosing: memory is almost never the limit; it is the choice of model that sets the speed, from 4.8 to 81.4 tokens per second.
What sets the speed: active parameters, not size
To generate each token, the chip must reread from memory the weights used for that token. With 273 GB/s of bandwidth, the maximum speed is simple to compute: 273 divided by the volume of weights read per token. A dense model reads all of its weights every time; an MoE model reads only a fraction, the “experts” chosen for that token.
This explains the table's paradox. The gpt-oss 120B file weighs 59 GB, but the model activates only 5.1 billion parameters per token: it generates at 58.0 tokens per second. Qwen3.8 27B, whose file is more than three times lighter (16 GB), is a dense model: it activates its 27 billion parameters on every token and generates at 11.8 tokens per second. The big model is nearly five times faster than the small one.
| Model | Active parameters | Theoretical ceiling | Measured | Share of ceiling |
|---|---|---|---|---|
| Qwen2.5-Coder 7B (dense) | 7.6B | 33.7 | 30.0 | 89% |
| Llama 3.3 70B (dense) | 70.6B | 6.4 | 4.8 | 74% |
| Qwen3.8 27B (dense) | 27.3B | 15.6 | 11.8 | 76% |
| Qwen3-Coder 30B-A3B (MoE) | 3.3B | 77.8 | 62.6 | 81% |
| gpt-oss 120B (MoE) | 5.1B | 98.7 | 58.0 | 59% |
| Qwen3.6 35B-A3B (MoE) | 3B | 141.1 | 66.0 | 47% |
The table covers six models, those whose number of active parameters is published or reported by llama.cpp. Dense models reach 74 to 89% of this ceiling: the GB10 gets almost everything out of its memory.
Across all thirteen models, the eight MoE models other than Qwen3.8-Flash-Next reach 47 to 81% of it, six of them between 52 and 62%; expert selection and side computations probably add a fixed time to each token. Qwen3.8-Flash-Next is left out of this calculation: it has 51 billion lookup-table parameters on top of its 125 billion, which makes the estimate unsuitable.
At comparable size, MoE models are still much faster. Hence our main advice: on a GB10 machine, for a large model, favor an MoE. A small dense model like Gemma 3 4B also generates very fast (80.8 tokens per second), but it is much smaller and serves other uses.
Related reading
Models of more than 100 billion parameters
This is what the 128 GB is for. NVIDIA advertises models of up to 200 billion parameters for the platform; our measurements confirm that promise, and a 235B compressed to 3 bits even fits beyond it. Five models of more than 100 billion parameters fit on the ATOM, all with at least 20 GB of margin at 30,000 tokens of context.
- gpt-oss 120B, the best compromise between speed and memory: 58.0 tokens per second, 61.7 GB used with its context, loaded in 10 seconds. OpenAI publishes it directly in MXFP4 format: it fits without additional compression.
- Qwen3.8-Flash-Next 125B, the fastest of the giant Qwens: 27.3 tokens per second with about 6 billion active parameters, 89.5 GB used in IQ4_XS. Its less compressed Q4_K_XL version fits too, with a narrow margin (see below).
- Qwen3.5 122B-A10B: 23.1 tokens per second, 74.5 GB; its ten billion active parameters put it behind gpt-oss.
- Nemotron-3 Super 120B-A12B (NVIDIA): 16.9 tokens per second, 80.4 GB. With twelve billion active parameters, it generates more slowly than gpt-oss, and it is the model whose generation holds up best as the context grows: 16.5 at 32,768 tokens, about 3% lower.
- Qwen3-235B-A22B, in a class of its own: It fits in Q2_K_XL, the strongest compression in the table (3 bits per parameter on average): 17.7 tokens per second, 90.5 GB. Compression this strong generally reduces the quality of the answers; we did not measure it.
Where memory runs out: around 100 to 105 GB of weights
The system sees 121.7 GB of memory: part of the 128 is reserved at boot. Once the machine was up, with nothing else in memory, 108 to 113 GB remained available depending on the moment. To find the ceiling, we loaded two heavier versions of the largest models, with the same 32,768-token context and a safeguard that cuts the load if free memory falls below 3 GB.
| Model | File | Memory used | Remaining margin | Generation (after 4,000 → 30,000 prompt tokens processed) | Fit |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next 125B (Q4_K_XL) | 103.7 GB | 107.0 GB | 6.0 GB | 24.9 → 21.9 tok/s | tight |
| Qwen3-235B-A22B (Q3_K_XL, 3.5 bits) | 97.0 GB | 104.7 GB | 8.2 GB | 13.9 → 10.4 tok/s | tight |
No model failed to load, but these two leave a narrow margin: hardly any room for a second model, a much longer context or a demanding application alongside. The practical limit is therefore around 100 to 105 GB of weights. It rules out, for example, the NVFP4 weights (a compressed NVIDIA format) of Qwen3.8-Flash-Next: about 135 GB according to the Kubesimplify blog (August 27, 2026), which states that two machines are then needed.
That same post already showed the model on a single machine in GGUF (llama.cpp's format), but only the most compressed version existed at the time. On the ATOM, the IQ4_XS and Q4_K_XL versions run, with 21 and 6 GB of margin.
The right size for comfortable use. Aim for about 90 GB at most including your context: that leaves around twenty GB for the system, a second small model or a longer context. Our thirteen models respect this guideline.
The cost of long context
A long document, a codebase or a conversation that keeps stretching: every token already in the context slows what follows. With 32,768 tokens in memory, generation speed drops by 3 to 47% depending on the model. We also measured the real time to process a 30,000-token document in one go, about fifty pages.
| Model | Processing 30,000 tokens | Generation afterward |
|---|---|---|
| Gemma 3 4B | 4.7 s | 60.8 tok/s |
| gpt-oss 20B | 8.1 s | 61.6 tok/s |
| Qwen2.5-Coder 7B | 12.7 s | 23.3 tok/s |
| Qwen3.6 35B-A3B | 14.2 s | 54.4 tok/s |
| Qwen3-Coder 30B-A3B | 17.4 s | 33.3 tok/s |
| gpt-oss 120B | 21.8 s | 42.8 tok/s |
| GLM-4.7-Flash | 32.7 s | 35.6 tok/s |
| Qwen3.8 27B | 39.4 s | 10.6 tok/s |
| Qwen3.8-Flash-Next 125B | 42.9 s | 23.0 tok/s |
| Qwen3.5 122B-A10B | 43.6 s | 21.2 tok/s |
| Nemotron-3 Super 120B-A12B | 55.4 s | 16.2 tok/s |
| Qwen3-235B-A22B | 1 min 33 s | 12.1 tok/s |
| Llama 3.3 70B | 1 min 44 s | 4.0 tok/s |
For most models, these times are longer than the “Prompt processing” column of the first table suggests. The probable reason: llama-server, the server used in practice, processes text in batches of 512 tokens by default, versus 2,048 in our llama-bench configuration.
Prompt processing is a strong point of the GB10: against the Mac compared further down, it processes gpt-oss 120B 31% faster. If you drop in a report of around fifty pages, the answer starts 22 seconds later with gpt-oss 120B, and 1 min 44 s later with a dense 70B. For document analysis and coding agents, this criterion weighs heavily.
Up to 16 users at once: what the machine can handle
With vLLM, we simulated 1, 8, then 16 simultaneous users. Each sends a request of 1,024 tokens and receives a response of 512 tokens; each point is measured three times with different requests, and we publish the median.
| Model | Users | Total | Per person | First word (average) | First word (slowest cases) |
|---|---|---|---|---|---|
| gpt-oss 120B | 1 | 35.6 tok/s | 36.4 tok/s | 0.34 s | 0.35 s |
| gpt-oss 120B | 8 | 113.9 tok/s | 14.5 tok/s | 0.97 s | 1.62 s |
| gpt-oss 120B | 16 | 160.3 tok/s | 10.2 tok/s | 1.07 s | 3.71 s |
| gpt-oss 20B | 1 | 49.1 tok/s | 50.0 tok/s | 0.16 s | 0.16 s |
| gpt-oss 20B | 8 | 205.9 tok/s | 26.5 tok/s | 0.49 s | 0.81 s |
| gpt-oss 20B | 16 | 322.4 tok/s | 20.7 tok/s | 0.52 s | 1.73 s |
From 1 to 16 users, total throughput is multiplied by 4.5 with gpt-oss 120B and by 6.6 with gpt-oss 20B: when several requests share the same reads of the model weights, the chip makes full use of its compute power.
At 16 people, each still sees the response written at 10 tokens per second with gpt-oss 120B, and at 21 with gpt-oss 20B. With the 120-billion model, the first word arrives in a little over a second on average, and in nearly 4 seconds in the slowest cases.
For a single person, however, vLLM is not the fastest: with no particular performance tuning, it generates at 36.4 tokens per second on gpt-oss 120B, against 54.4 for llama.cpp on long responses. Its logs show that on the GB10 it picks the Marlin compute kernel for the MXFP4 format, which probably explains the gap. vLLM makes sense as soon as several people or several agents share the machine.
Alone at the machine: Ollama or llama.cpp?
Ollama is the simplest way to get started, and it works on the GB10 with no tuning. On gpt-oss, though, it is not the fastest. On long responses, version 0.34.1 generates at 42.3 tokens per second on gpt-oss 120B, against 54.4 for llama.cpp in the same MXFP4 format, which is 22% less. On gpt-oss 20B: 58.8 against 78.8, which is 25% less.
On Qwen models, it is the opposite. On Qwen3.8 27B, Ollama generates between 22.9 and 30.5 tokens per second depending on the run, against 11.8 for llama.cpp with default settings; on Qwen3.6 35B-A3B, between 90.5 and 94.9, against 66.0. The probable reason: Ollama enables speculative decoding by default there, which proposes several tokens ahead that the model validates in one go (the draft_num_predict setting appears in the configuration of these models).
Our test points the same way. On llama-server (six texts of 400 tokens, prose and code), llama.cpp's speculative decoding (MTP, which predicts several tokens at once) takes Qwen3.8 27B from 11.7 to 21.5 tokens per second on prose, and from 11.6 to 27.2 on code.
Ollama opens long contexts by default. With no adjustment from us, Ollama 0.34.1 opened contexts of 131,072 tokens for gpt-oss and 262,144 for Qwen and Gemma, even though its documentation says 4,096 by default. The memory it reports stays close to our measurements. To keep room for a second model, reduce the context with the OLLAMA_CONTEXT_LENGTH variable.
In practice: Ollama to get started and for Qwen models, which it speeds up out of the box; llama.cpp to get the most out of gpt-oss, or Qwen3.8 with its speculative decoding turned on. The best choice depends on settings more than on the machine.
Versus the MacBook Pro M5 Max 128 GB
A reader measured their MacBook Pro M5 Max 128 GB (40-core GPU) on September 16 with Ollama 0.34.1, in a single run per model; the readings are published in our dedicated guide. We replayed the same series on the ATOM, in two runs: same Ollama version, same models, same prompt.
| Measurement | ATOM, 1st run | ATOM, 2nd run | MacBook Pro M5 Max | Difference |
|---|---|---|---|---|
| Generation, gemma4:12b | 45.9 | 54.1 | 58.1 | Mac +7 to +27% |
| Generation, qwen3.8:27b | 30.5 | 22.9 | 36.6 | Mac +20 to +59% |
| Generation, gpt-oss:20b | 58.1 | 58.8 | 113.4 | Mac +93 to +95% |
| Generation, gpt-oss:120b | 42.2 | 42.3 | 79.1 | Mac +87% |
| Prompt processing, gpt-oss:120b | 1,816 | — | 1,388 | ATOM +31% |
The Mac generates faster on all four models, nearly twice as fast on the two gpt-oss models, whose runs agree: its chip has 614 GB/s of memory bandwidth, more than double the GB10's. On gemma4 and qwen3.8, the gap varies from one run to the next; speculative decoding, whose gain depends on the generated text, is one possible explanation.
The ATOM processes prompts faster, probably thanks to the compute power of its GPU, on texts that are close but not identical (16,850 and 16,689 tokens). With llama.cpp, the generation gap narrows: 54.4 tokens per second on gpt-oss 120B, against 79.1 for the Mac under Ollama.
Related reading
Who the ATOM is the right choice for
What follows comes from our measurements on the ATOM.
- The ATOM is the right choice if you want large models at home: Five models of more than 100 billion parameters fit with room to spare. To our knowledge, no consumer graphics card comes close to these 128 GB.
- … if you work on long documents or code: 30,000 tokens processed in 22 seconds with gpt-oss 120B.
- … if several people or agents share it: 160 tokens per second in total for 16 users on gpt-oss 120B.
- … if you want the NVIDIA ecosystem: CUDA 13, vLLM in NVIDIA's container and llama.cpp compiled for the GB10 chip all worked for us under DGX OS 7.5.0.
- If you work alone and mainly care about generation speed: Compare with the MacBook Pro M5 Max: thanks to its 614 GB/s, it generates faster on gpt-oss, while the ATOM processes a long document 31% faster on gpt-oss 120B. Compare prices too.
- If your models stay under 35 GB: The ATOM's 64 GB version, announced for October 23, 2026, should be enough (a calculation drawn from our measurements, not measured on that version).
Related reading
64 or 128 GB?
On October 2, 2026, NVIDIA announced 64 GB DGX Spark machines for models of up to 100 billion parameters, available on October 23 from Acer, ASUS, Dell, GIGABYTE, HP and MSI, starting at $4,999. GIGABYTE confirmed a 64 GB AI TOP ATOM with the same design on October 5. We did not measure it.
Our readings allow a calculation. Models up to Qwen3-Coder 30B-A3B use less than 35 GB with 32,768 tokens of context and would fit; Llama 3.3 70B (51 GB) would be at the limit. gpt-oss 120B (62 GB) and larger models would not fit. At equal bandwidth, which remains to be verified, speeds should be close.
For a model of more than 100 billion parameters on a single machine, the 128 GB version is the one to get. According to NVIDIA, two linked 64 GB machines also pool their memory; we did not test that.
Prices checked
On October 8, 2026, in France, the 4 TB PCIe 4.0 version (ATAGB10-9001, same GB10 chip and same 128 GB) was listed at €6,346.27 including VAT on the official AORUS store, sold out; it was in stock on Amazon.fr, sold by Amazon UK (one unit). The tested version, 4 TB PCIe 5.0 (ATAGB10-9000), was €7,999.95 at LDLC and at Materiel.net, out of stock.
Prices and stock for these machines change quickly: check them before you buy.
Our verdict
The ATOM is an excellent choice for running a 120-billion-parameter model at home, sharing it with a team or processing long documents quickly. It gave us the platform's reference performance, with no thermal throttling flagged in our readings, including during one hour of continuous load at 16 users. If budget matters, its PCIe 4.0 version keeps the same chip and the same memory.
Related reading
What we did not measure
- Answer quality: This guide measures what fits and how fast, not how good each model is.
- Temperatures, noise and power consumption: The next guide in the series details temperatures under long load; power consumption has its own guide, read from the chip (GPU only). For noise, we cite Hardware & Co's measurements.
- Contexts beyond 32,768 tokens: Several models accept much longer contexts; this guide stops at 32,768 tokens, and the series' 70B tutorial goes up to 131,072 tokens for Llama 3.3 70B and gpt-oss 120B.
- Other GB10 machines and the 64 GB version: Our figures line up with those published for NVIDIA's DGX Spark, but we only measured the 128 GB ATOM.
Updates and corrections: this page will be revised if a new version of DGX OS, llama.cpp or Ollama changes these results; each correction will be dated.
Related reading
- The data behind our measurements, table by table (CSV)
- Our measurement methodology and its dated addenda
Sources
- NVIDIA, DGX Spark platform specifications (GB10, 273 GB/s, models up to 200 billion parameters)
- llama.cpp, reference table published for the DGX Spark (build 7941)
- Ollama, published measurements for the DGX Spark (October 23, 2025)
- OpenAI, gpt-oss 120B model card (5.1 billion active parameters, MXFP4 format)
- Kubesimplify, Qwen3.8-Flash-Next on DGX Spark (August 27, 2026)
- NVIDIA, announcement of the 64 GB DGX Spark (October 2, 2026)
- GIGABYTE, announcement of the 64 GB AI TOP ATOM (October 5, 2026)
- Apple, MacBook Pro M5 Max specifications (614 GB/s)
- Price check of October 8, 2026: official AORUS store (ATAGB10-9001)
- Price check of October 8, 2026: LDLC (ATAGB10-9000)
- Price check of October 8, 2026: Amazon.fr (ATAGB10-9001)
- Price check of October 8, 2026: Materiel.net (ATAGB10-9000)
FAQ
What is the largest model you can run on a 128 GB GB10 machine?
On our AI TOP ATOM, the largest one tested is Qwen3-235B-A22B: 90.5 GB used in Q2_K_XL (3 bits per parameter) and 21 GB of margin at 30,000 tokens of context. Its Q3_K_XL version fits too, with 8 GB of margin. We did not measure the quality of the answers: at 3 bits, a model strays further from its original version than at 4 or 5 bits, but that does not tell you which one answers best.
Is a 70B model usable day to day on a GB10?
It fits without difficulty (51 GB used) and generates 4.8 tokens per second; it processes 30,000 tokens in 1 min 44 s. It is ideal for batch processing: in NVFP4 under vLLM, eight simultaneous requests total 37.7 tokens per second, and each response, once started, is generated at 5.0 tokens per second. A small draft model brings it to 12 tokens per second (see our 70B tutorial). For conversation, gpt-oss 120B is much more pleasant.
Do this guide's speeds apply to an NVIDIA DGX Spark or another brand?
Very probably, within a few percent, but we only measured the ATOM. To our knowledge, 128 GB GB10 machines share the same chip and the same memory. Replaying two published references for the DGX Spark, with llama.cpp and with Ollama, our prompt processing speeds are within 2.3% and our generation speeds are at most 4.3% lower.
Should I choose the 64 GB or the 128 GB version?
If your models use less than 35 GB with their context, like gpt-oss 20B or Qwen3.6 35B-A3B in our measurements, the 64 GB version should be enough; that is a calculation, we did not measure it. For a model of more than 100 billion parameters on a single machine, like gpt-oss 120B (62 GB), the 128 GB version is required.
Is Ollama the best choice on a GB10 machine?
To get started, yes: it works with no tuning. On gpt-oss 120B, though, llama.cpp generates faster (54.4 versus 42.3 tokens per second on long responses). On Qwen models, Ollama wins with its default settings, probably thanks to the speculative decoding it enables; llama.cpp gets close when you turn on its own. The best choice therefore depends mostly on settings.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.