NVIDIA DGX Spark: price, specs, and verdict for local AI (2026)
The compact AI workstation from NVIDIA, with 128 GB of unified memory. Let’s look at what matters for running LLMs locally : memory, bandwidth, speed, price—and who it's (or isn't) the right purchase for.
Updated on 09/10/2026
Where to buy it
The button configuration is the one offered for purchase. On mini PCs, keep some memory available for the system and check the inference engine; macOS/MLX and CUDA are not interchangeable.
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence these independently determined recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
Technical specifications
- Memory128 GB unified (LPDDR5x)
- Bandwidth273 GB/s
- Compute≈ 1 PFLOP FP4 — Blackwell GPU + 20-core Arm Grace CPU
- Power consumption140 W
- Indicative price≈ 6,100 € at Materiel.net (DGX Spark PNY at 6,099.95 € on 28/09/2026, lead time of more than 15 days; the MSI EdgeXpert, with the same GB10 chip and 128 GB, is 6,599.95 € with delivery within 7 days)
- Largest modelNVIDIA positions it at “up to 200B” since the CES 2026 software update; in practice ~70B in comfortable Q4, limited by bandwidth (273 GB/s). Two chained units run a Qwen 3.5-122B at ~40 t/s
Who is it for?
Advantages and limitations
✓ Strengths
- 128 GB of coherent unified memory
- Complete CUDA stack, locally
- Compact form factor, 140 W
- Perfect for on-premises / air-gapped / compliance
- CES 2026 software update: up to +2.5× faster inference (TensorRT-LLM, speculative decoding)
- Expandable: 2 chained units = a 122B at ~40 t/s
✗ Limitations
- Modest 273 GB/s bandwidth (a Mac Studio delivers ~2×)
- High price relative to raw performance
- Often very limited storage
- Difficult to recommend outside the CUDA / on-prem ecosystem
Choices and limitations — review of September 12, 2026
The right criterion: your model and your environment
DGX Spark offers 128 GB of unified memory in a NVIDIA environment. An advertised model capacity guarantees neither the desired context nor interactive throughput. Check the quantization, runtime, number of users, and whether the measurement concerns one or two units.
Sources: NVIDIA manufacturer spec sheet et external measurements and raw data on a pair of DGX Spark systems. No in-house DGX Spark benchmark is claimed in this sheet; do not apply the results from a pair to a single unit.
Return to hardware selection · Install your first local AI · Examine our protocol on a different machine
Dense or MoE: what flies and what crawls on the DGX Spark
Memory bandwidth (273 GB/s) is the machine's real bottleneck. With each generated token, a model dense rereads all of its weights; a MoE (Mixture of Experts) reads only its active experts. On the DGX Spark, the difference is spectacular:
| Model | Architecture | Measured speed |
|---|---|---|
| gpt-oss-120B (MXFP4/FP8) | MoE — ~5B active/token | ≈ 40–58 tok/s depending on the engine (llama.cpp, SGLang) |
| Qwen3-Coder 30B-A3B | MoE — 3B active/token | Very smooth, ideal for local coding |
| Llama 3.1 70B (FP8) | Dense | ≈ 2.7 tok/s (LMSYS measurement)—uncomfortable for daily use |
Practical conclusion: on this machine, prioritize MoE models — this is where the 128 GB really matter. Our DGX Spark guide for local AI details the settings, and the catalog filters models by VRAM.
Compared with 128 GB alternatives
Three desktop machines can now fit a very large model in unified memory:
| Machine | Memory | Bandwidth | Indicative price | Ecosystem |
|---|---|---|---|---|
| NVIDIA DGX Spark | 128 GB | 273 GB/s | ≈ 6 100–6 600 € (Materiel.net, 28/09/2026) | Full CUDA |
| Ryzen AI Max+ 395 Mini PC | up to 128 GB | 256 GB/s | ≈ €1,900–3,900 depending on configuration | ROCm / Vulkan |
| Mac Studio M5 Max | 36-128 GB | 460–614 GB/s | starting at €2,999 (36 GB) | MLX / Metal |
At the same budget, Strix Halo runs the same MoE models for 2 to 4× less; the Mac is significantly faster in bandwidth. The DGX Spark retains one unique advantage: the full CUDA stack within 140 W. Detailed comparison: Strix Halo vs DGX Spark.
DGX Spark in the enterprise: total cost, break-even point, and a two-unit cluster
Content taken from the “DGX and Spark: Cost Analysis for IT Departments” guide (merged into this page on 08/09/2026). The official price rose from $3 999 to $4 699 on 27 February 2026 (+17.5%), NVIDIA citing supply constraints affecting LPDDR5X; in France, the price observed at third-party retailers remains significantly higher (see the price section above).
- Break-even point : as soon as a recurring cloud inference bill exceeds ~€300 to €500 per month, buying becomes rational over 12 to 18 months—even before factoring in sovereignty (GDPR, medical, or legal data that never leaves the building). This is a one-time CAPEX amortized over 3-4 years versus a continuous cloud bill.
- What 128 GB really enables : a 70B model in Q4_K_M occupies ≈ 40 GB, and a heavily quantized 120B model occupies ≈ 70–80 GB; there is still room for context and the KV cache, whereas a 24 GB card tops out at around 32B.
- CES 2026 Software Award : up to ×2.5 on the same machine via TensorRT-LLM and speculative decoding—maximum gains on predictable outputs (code, JSON, structured text), with more modest gains on creative text.
- Two chained units (ConnectX): 256 GB aggregated, Qwen 3.5-122B at ~40 tokens/s; budget ≈ $9,400 excl. tax. Chaining doubles the memory accessible, but not necessarily the throughput of a model that already fit on a single machine.
- Selection criterion : CUDA + clustering → DGX Spark; tight budget → Ryzen AI Max+ 395 mini PC (128 GB, less mature ROCm ecosystem); bandwidth and versatility → Mac Studio (> 800 GB/s, MLX/Metal).
For a company that prefers to delegate integration rather than build this type of machine itself, IAPRO offers a turnkey hardware + AI package for SMBs : server selection, model deployment, and team training.
Verdict
Frequently asked questions
Can the NVIDIA DGX Spark run a 70B LLM locally?
NVIDIA positions it as “up to 200B” since the CES 2026 software update; in practice, ~70B is comfortable in Q4, limited by bandwidth (273 GB/s). Two chained units run a Qwen 3.5-122B at ~40 t/s.
How much does the NVIDIA DGX Spark cost?
Expect to pay ≈ 6 100 € at Materiel.net (DGX Spark PNY at 6 099,95 € on 28/09/2026, with a lead time of more than 15 days; the MSI EdgeXpert, with the same GB10 chip and 128 GB, costs 6 599,95 € with delivery within 7 days) (officially 4 699 $ (3 999 $ recorded on 27 February 2026)). Prices change quickly—check the current price through the purchase links.
Who is the NVIDIA DGX Spark for?
Developers committed to CUDA, local fine-tuning, and on-premises / air-gapped environments.
Where to buy the DGX Spark in France?
Materiel.net is the benchmark: the PNY DGX Spark was priced there at 6 099,95 € on 28/09/2026, with a wait of more than 15 days. Elsewhere, it is mostly sold by third-party sellers at higher, volatile prices. Serious alternative: partners’ GB10 machines (Asus Ascent GX10, Dell, HP, Lenovo…), with the same chip and memory, often better distributed—the full list is in our DGX Spark guide.
DGX Spark or RTX 5090 for local AI?
Two philosophies: the RTX 5090 offers 32 GB of GDDR7 at ~1.8 TB/s — any model that fits within 32 GB will run much faster on it. The DGX Spark offers 4× the capacity (128 GB), but at 273 GB/s: it can fit models that the 5090 cannot load, serving them more slowly. If your target models are under 32 GB, choose the GPU; see choosing a GPU for local AI.
Should you wait for the “RTX Spark” PCs instead of buying a DGX Spark?
The RTX Spark models, based on the GB10 chip and offering up to 128 GB of unified memory, are confirmed for fall 2026 from Asus, Dell, HP, Lenovo, MSI, and Microsoft—eight announced machines, including laptops. But as of August 31, 2026, no price has been published and no preorders are open. If you need CUDA immediately, the DGX Spark remains the only shippable option; otherwise, waiting until fall will significantly expand the selection.
Go further
Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.