Home›AI hardware›NVIDIA DGX Spark

NVIDIA DGX Spark: price, specs, and verdict for local AI (2026)

The compact AI workstation from NVIDIA, with 128 GB of unified memory. Let’s look at what matters for running LLMs locally : memory, bandwidth, speed, price—and who it's (or isn't) the right purchase for.

Updated on 09/10/2026

Where to buy it

NVIDIA DGX Spark
AmazonBOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

The button configuration is the one offered for purchase. On mini PCs, keep some memory available for the system and check the inference engine; macOS/MLX and CUDA are not interchangeable.

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence these independently determined recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Technical specifications

Who is it for?

Developers committed to CUDA, local fine-tuning, and on-premises / air-gapped environments.

Advantages and limitations

✓ Strengths

  • 128 GB of coherent unified memory
  • Complete CUDA stack, locally
  • Compact form factor, 140 W
  • Perfect for on-premises / air-gapped / compliance
  • CES 2026 software update: up to +2.5× faster inference (TensorRT-LLM, speculative decoding)
  • Expandable: 2 chained units = a 122B at ~40 t/s

✗ Limitations

  • Modest 273 GB/s bandwidth (a Mac Studio delivers ~2×)
  • High price relative to raw performance
  • Often very limited storage
  • Difficult to recommend outside the CUDA / on-prem ecosystem

Choices and limitations — review of September 12, 2026

The right criterion: your model and your environment

DGX Spark offers 128 GB of unified memory in a NVIDIA environment. An advertised model capacity guarantees neither the desired context nor interactive throughput. Check the quantization, runtime, number of users, and whether the measurement concerns one or two units.

Sources: NVIDIA manufacturer spec sheet et external measurements and raw data on a pair of DGX Spark systems. No in-house DGX Spark benchmark is claimed in this sheet; do not apply the results from a pair to a single unit.

Return to hardware selection · Install your first local AI · Examine our protocol on a different machine

Dense or MoE: what flies and what crawls on the DGX Spark

Memory bandwidth (273 GB/s) is the machine's real bottleneck. With each generated token, a model dense rereads all of its weights; a MoE (Mixture of Experts) reads only its active experts. On the DGX Spark, the difference is spectacular:

ModelArchitectureMeasured speed
gpt-oss-120B (MXFP4/FP8)MoE — ~5B active/token≈ 40–58 tok/s depending on the engine (llama.cpp, SGLang)
Qwen3-Coder 30B-A3BMoE — 3B active/tokenVery smooth, ideal for local coding
Llama 3.1 70B (FP8)Dense≈ 2.7 tok/s (LMSYS measurement)—uncomfortable for daily use

Practical conclusion: on this machine, prioritize MoE models — this is where the 128 GB really matter. Our DGX Spark guide for local AI details the settings, and the catalog filters models by VRAM.

Compared with 128 GB alternatives

Three desktop machines can now fit a very large model in unified memory:

MachineMemoryBandwidthIndicative priceEcosystem
NVIDIA DGX Spark128 GB273 GB/s≈ 6 100–6 600 € (Materiel.net, 28/09/2026)Full CUDA
Ryzen AI Max+ 395 Mini PCup to 128 GB256 GB/s≈ €1,900–3,900 depending on configurationROCm / Vulkan
Mac Studio M5 Max36-128 GB460–614 GB/sstarting at €2,999 (36 GB)MLX / Metal

At the same budget, Strix Halo runs the same MoE models for 2 to 4× less; the Mac is significantly faster in bandwidth. The DGX Spark retains one unique advantage: the full CUDA stack within 140 W. Detailed comparison: Strix Halo vs DGX Spark.

DGX Spark in the enterprise: total cost, break-even point, and a two-unit cluster

Content taken from the “DGX and Spark: Cost Analysis for IT Departments” guide (merged into this page on 08/09/2026). The official price rose from $3 999 to $4 699 on 27 February 2026 (+17.5%), NVIDIA citing supply constraints affecting LPDDR5X; in France, the price observed at third-party retailers remains significantly higher (see the price section above).

For a company that prefers to delegate integration rather than build this type of machine itself, IAPRO offers a turnkey hardware + AI package for SMBs : server selection, model deployment, and team training.

Verdict

Best reserved for developers committed to CUDA or for on-premises / air-gapped needs. For large models without that constraint, also compare a Strix Halo mini PC or a Mac. On AMD, check whether your tools support Vulkan or ROCm: 128 GB of shared memory is not a substitute for a CUDA environment. Note: the same GB10 chip arrives in fall 2026 in «RTX Spark» laptops and desktops (Dell, Asus, HP, Lenovo…) — the 128 GB unified-memory offering will expand.

Frequently asked questions

Can the NVIDIA DGX Spark run a 70B LLM locally?

NVIDIA positions it as “up to 200B” since the CES 2026 software update; in practice, ~70B is comfortable in Q4, limited by bandwidth (273 GB/s). Two chained units run a Qwen 3.5-122B at ~40 t/s.

How much does the NVIDIA DGX Spark cost?

Expect to pay ≈ 6 100 € at Materiel.net (DGX Spark PNY at 6 099,95 € on 28/09/2026, with a lead time of more than 15 days; the MSI EdgeXpert, with the same GB10 chip and 128 GB, costs 6 599,95 € with delivery within 7 days) (officially 4 699 $ (3 999 $ recorded on 27 February 2026)). Prices change quickly—check the current price through the purchase links.

Who is the NVIDIA DGX Spark for?

Developers committed to CUDA, local fine-tuning, and on-premises / air-gapped environments.

Where to buy the DGX Spark in France?

Materiel.net is the benchmark: the PNY DGX Spark was priced there at 6 099,95 € on 28/09/2026, with a wait of more than 15 days. Elsewhere, it is mostly sold by third-party sellers at higher, volatile prices. Serious alternative: partners’ GB10 machines (Asus Ascent GX10, Dell, HP, Lenovo…), with the same chip and memory, often better distributed—the full list is in our DGX Spark guide.

DGX Spark or RTX 5090 for local AI?

Two philosophies: the RTX 5090 offers 32 GB of GDDR7 at ~1.8 TB/s — any model that fits within 32 GB will run much faster on it. The DGX Spark offers 4× the capacity (128 GB), but at 273 GB/s: it can fit models that the 5090 cannot load, serving them more slowly. If your target models are under 32 GB, choose the GPU; see choosing a GPU for local AI.

Should you wait for the “RTX Spark” PCs instead of buying a DGX Spark?

The RTX Spark models, based on the GB10 chip and offering up to 128 GB of unified memory, are confirmed for fall 2026 from Asus, Dell, HP, Lenovo, MSI, and Microsoft—eight announced machines, including laptops. But as of August 31, 2026, no price has been published and no preorders are open. If you need CUDA immediately, the DGX Spark remains the only shippable option; otherwise, waiting until fall will significantly expand the selection.

Go further

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.