NVIDIA DGX Spark: review for local AI
NVIDIA's compact AI workstation with 128 GB of unified memory. We look at what matters for running LLMs locally — memory, bandwidth, speed, price — and who it is (or is not) the right buy for.
Where to buy
The button names the configuration offered. A mini PC is a complete PC alternative: check usable memory and software support. It does not replace macOS/MLX or CUDA.
As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. This does not influence our independent recommendations.
Specs
| Memory | 128 GB unified (LPDDR5x) |
| Bandwidth | 273 GB/s |
| Compute | ~1 PFLOP FP4 — Blackwell GPU + 20-core Grace Arm CPU |
| Power | 140 W |
| Price | ~$4,000–4,700 |
| Largest model | Models up to ~70B at Q4, but bandwidth-limited (273 GB/s) |
Who is it for?
CUDA-committed developers, local fine-tuning, on-premise / air-gapped setups.
Pros and cons
Pros:
- 128 GB of coherent unified memory
- Full CUDA stack, locally
- Compact, 140 W
- Great for on-premise / air-gapped / compliance
Cons:
- Modest 273 GB/s bandwidth (a Mac Studio does ~2x)
- Pricey for the raw performance
- Often very limited stock
- Hard to recommend outside the CUDA / on-prem use case
Verdict
Reserve it for developers committed to CUDA or for on-premise / air-gapped needs. To run a 70B without that constraint, a Mac Studio or a Strix Halo mini PC usually offer a better price/performance ratio.
FAQ
Can the NVIDIA DGX Spark run a 70B LLM locally?
Models up to ~70B at Q4, but bandwidth-limited (273 GB/s).
How much does the NVIDIA DGX Spark cost?
Expect ~$4,000–4,700. Prices move fast — check today's price via the buy links.
Who is the NVIDIA DGX Spark for?
CUDA-committed developers, local fine-tuning, on-premise / air-gapped setups.