NVIDIA DGX Spark: review for local AI
NVIDIA's compact AI workstation with 128 GB of unified memory. We look at what matters for running LLMs locally — memory, bandwidth, speed, price — and who it is (or is not) the right buy for.
Where to buy
As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. This does not influence our independent recommendations.
Specs
| Memory | 128 GB unified (LPDDR5x) |
| Bandwidth | 273 GB/s |
| Compute | ~1 PFLOP FP4 — Blackwell GPU + 20-core Grace Arm CPU |
| Power | 140 W |
| Price | ~$4,000–4,700 |
| Largest model | Models up to ~70B at Q4, but bandwidth-limited (273 GB/s) |
Who is it for?
CUDA-committed developers, local fine-tuning, on-premise / air-gapped setups.
Pros and cons
Pros:
- 128 GB of coherent unified memory
- Full CUDA stack, locally
- Compact, 140 W
- Great for on-premise / air-gapped / compliance
Cons:
- Modest 273 GB/s bandwidth (a Mac Studio does ~2x)
- Pricey for the raw performance
- Often very limited stock
- Hard to recommend outside the CUDA / on-prem use case
Verdict
Reserve it for developers committed to CUDA or for on-premise / air-gapped needs. To run a 70B without that constraint, a Mac Studio or a Strix Halo mini PC usually offer a better price/performance ratio.
FAQ
Can the NVIDIA DGX Spark run a 70B LLM locally?
Models up to ~70B at Q4, but bandwidth-limited (273 GB/s).
How much does the NVIDIA DGX Spark cost?
Expect ~$4,000–4,700. Prices move fast — check today's price via the buy links.
Who is the NVIDIA DGX Spark for?
CUDA-committed developers, local fine-tuning, on-premise / air-gapped setups.