NVIDIA DGX Spark: the 128 GB AI mini-PC versus GPU
The DGX Spark is NVIDIA's bet on local desktop AI: a Grace Blackwell chip, 128 GB of unified memory, and a form factor that fits on a corner of your desk. On paper, it runs models that no consumer GPU can load on its own. This guide takes a practical look at the dgx spark llm: what fits in memory, real inference speeds, and a head-to-head comparison with a RTX 5090 and a Mac Studio to determine whether this mini PC deserves a place in your setup.
Choosing a machine? Our picks by budget → · Our spec sheet NVIDIA DGX Spark →
Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).
Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What the DGX Spark is
The DGX Spark is a compact computer built around the NVIDIA GB10 Grace Blackwell superchip. It combines a Blackwell GPU and a Grace ARM CPU (20 cores) on the same chip, connected to a single pool of 128 GB of LPDDR5X memory shared between the two. This unified-memory architecture—the same philosophy as Apple Silicon—is what makes the machine so compelling for local AI.
Where a conventional graphics card isolates its VRAM (24 or 32 GB maximum for consumer hardware), the DGX Spark lets the GPU address all 128 GB directly. In practice, the constraint that blocks everyone on GPUs—“does the model fit in VRAM?”—almost disappears. The limiting factor becomes memory bandwidth, not capacity.
- Chip
- NVIDIA GB10 Grace Blackwell (Blackwell GPU + 20-core ARM CPU)
- Memory
- 128 GB unified LPDDR5X, shared by the CPU/GPU
- Bandwidth
- ~273 GB/s (LPDDR5X), the real bottleneck
- Network
- ConnectX (200 GbE) to link two Sparks and target very large models
- Format
- Desktop mini PC, ~150 W, quiet compared with a multi-GPU tower
#128 GB unified memory: what it unlocks
For context, here is the approximate VRAM required by a model in Q4_K_M quantization, the most common choice for local inference: a 7B fits in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B in ~40 GB. On a RTX 5090 with its 32 GB, a 70B Q4 therefore does not fit entirely: part of it must be offloaded to the CPU, which causes a dramatic drop in speed.
With 128 GB of unified memory, the DGX Spark easily loads a 70B model in Q4, Q5, or even Q8, leaves room for a long context (32k tokens and more), and opens the door to 100B+ parameter models with aggressive quantization. This is the class of model that, until now, required either a high-end Mac Studio or an expensive, power-hungry multi-GPU rig.
- 70B Q4 (~40 GB)
- Comfortable, with plenty of room for context
- 70B Q8 (~75 GB)
- Possible—near-FP16 quality on a large model
- 120B-type MoE in Q4
- Fits in memory; bandwidth determines speed
- Multiple models loaded
- One 32B + one 14B in parallel for routing or multi-agent workloads
#Which models it runs
The DGX Spark's strength is the 30B–120B range, exactly where consumer GPUs stall due to insufficient VRAM. Here are the typical uses by model family.
- Qwen 3.8 27B
- The flagship generalist of 2026 (vision, 262k ctx); on this machine it runs at full-quality Q8 with plenty of context headroom
- Qwen 3.6 35B-A3B / Qwen3-Coder 30B
- Fast MoEs for reasoning and coding, fully resident in memory even in Q8
- GLM 4.7 Flash (MoE 30B-A3B, MIT)
- Excellent for agents, loaded at full quality without juggling
- Large 100B+ quantized MoE
- Mixture-of-Experts architectures fully benefit from the large unified memory
- Multimodal models
- Gemma 4 and Qwen 3.8 (vision + text), memory-hungry, run without compromise
For small models (7B–14B), the DGX Spark obviously works, but it's overkill: a RTX 4070 or an M4 Mac mini runs them faster for far less money. The machine only makes sense if you regularly target large models.
#Real-world inference performance
Two figures matter in local inference: generation speed (tokens/sec, what scrolls across the screen) and prefill time (the delay before the first token, which depends on prompt length). Memory bandwidth primarily drives generation.
On the DGX Spark, a 70B model in Q4 typically generates around ten tokens/sec—enough for a comfortable assistant experience where you read as output arrives, but less suitable for high-throughput batch workloads. Smaller models (14B–32B) go significantly faster. This profile is consistent with ~273 GB/s of bandwidth: that's the wall separating the DGX Spark from much faster GDDR7 cards in this respect.
#DGX Spark vs RTX 5090
This is the most revealing comparison, and it pits two philosophies against each other. The RTX 5090 offers 32 GB of GDDR7 VRAM with ~1792 GB/s of bandwidth—about six to seven times faster than the DGX Spark's memory. For anything that fits within its 32 GB, the 5090 crushes the Spark in tokens/sec.
But the 5090 hits a hard wall: beyond ~32 GB, it has to offload to the CPU and performance collapses. A 70B Q4 (~40 GB) doesn’t fit entirely in it. The DGX Spark, by contrast, handles that same 70B without flinching—more slowly, but as a single piece. The choice comes down to one question: does your workload fit within 32 GB?
- You stay at ≤ 32 GB (up to ~32B Q4)
- RTX 5090: much faster and versatile (gaming, rendering, CUDA)
- You are targeting 70B and above
- DGX Spark: the model fits where the 5090 gives up
- Power consumption / noise
- DGX Spark: ~150 W and quiet, versus a hot, noisy 5090 tower
- Complete CUDA ecosystem
- Both; the Spark adds the DGX/NVIDIA stack end to end
#DGX Spark vs. Mac Studio
The DGX Spark's real rival isn't a graphics card; it's the Mac Studio. Both share the same idea—massive unified memory in a compact, efficient enclosure—and target the same audience: developers who want large models locally without building a rig.
The Mac Studio (M-Ultra) supports up to 512 GB of unified memory with often higher bandwidth on high-end configurations, making it formidable for very large models. Its advantage: macOS, Metal, and a mature ecosystem (Ollama, LM Studio, MLX). The DGX Spark's advantage: native CUDA. If your pipeline depends on CUDA libraries, TensorRT, or NVIDIA tools, the Spark natively speaks the same language as your production servers.
- Maximum memory capacity
- Mac Studio up to 512 GB; DGX Spark 128 GB (expandable by chaining 2 units)
- Software ecosystem
- Mac: mature Metal/MLX; Spark: native CUDA, continuity with NVIDIA servers
- AI code portability
- The Spark runs code written for NVIDIA cloud/datacenter GPUs without friction
- Office use
- The Mac Studio is a complete workstation; the Spark is a dedicated AI station
#Install Ollama on it
The DGX Spark runs on a Linux base (DGX OS, derived from Ubuntu) with the CUDA stack preinstalled. Ollama works there like on any Linux machine NVIDIA: the daemon detects the GPU and serves models on its usual port.
- 01Install OllamaOfficial one-command script. It configures the systemd service and automatically detects the Blackwell GPU through CUDA.
- 02Check GPU detectionConfirm that the GPU is detected before launching a large model; otherwise, inference would fall back to the CPU.
- 03Pull a large modelDownload a large 2026 MoE at full quality (a Qwen3-Coder 30B in Q8, or the flagship Qwen 3.8 27B with its 262k context) to use the available memory—this is exactly what the machine is designed for.
- 04Launch and connect an interfaceThe daemon listens on http://localhost:11434 by default; point Open WebUI or LM Studio to it.
The Ollama daemon exposes a compatible API on port 11434. To make the most of your memory, favor models that standard GPUs cannot load on their own, and increase the context window since you have room for it. Note that Qwen 3.8 27B tends to overthink with its default reasoning setting: switch it to low mode for more direct answers.
#One chip, eight machines: certified clones
NVIDIA turned GB10 into a platform: besides its DGX Spark Founders Edition, seven manufacturers sell the same certified base under their own brands. Same chip, same 128 GB—the differences come down to storage, chassis, price, and sales channel.
| Machine | Manufacturer | Key takeaway |
|---|---|---|
| DGX Spark Founders Edition | NVIDIA | The benchmark, ~$3,999 |
| Ascent GX10 | Asus | The best price point (~$2,999) |
| Veriton GN100 | Acer | Single SKU with 4 TB of storage |
| AI TOP ATOM | Gigabyte | Comes with the AI TOP software suite |
| EdgeXpert MS-C931 | MSI | Industrial/edge channel |
| Pro Max GB10 | Dell | Enterprise channel |
| ThinkStation PGX | Lenovo | Workstation channel |
| ZGX Nano | HP | Workstation channel |
A detail that changes everything for large models: the integrated ConnectX port lets you link two units and run 400B-class models locally. No other consumer machine offers this.
#Frequently asked questions
What exactly is the DGX Spark's GB10?+
What is the difference between the DGX Spark and its Asus, Acer, or Gigabyte clones?+
Can you really run a 400B model on two machines?+
DGX Spark or Strix Halo machine?+
Who is a GB10 machine for?+
#Who this format makes sense for
The DGX Spark is not a consumer machine. Its value is concentrated among specific users for whom the VRAM constraint is the real day-to-day problem.
- AI developers on CUDA
- Locally prototype code intended for NVIDIA production GPUs, without porting surprises
- Researchers and R&D teams
- Light fine-tuning and inference of large models without reserving cloud GPU time
- 70B+ power users
- Those who constantly run into the 24–32 GB limit of a consumer graphics card
- Noise- and power-sensitive offices
- A quiet, efficient workstation rather than a multi-GPU rig
Conversely, if you game, mainly run models ≤ 32B, or have a tight budget, a RTX 5090 (or even a 4070/4080) or an M4 Mac mini will offer much better performance per dollar. The DGX Spark makes sense when massive memory capacity matters more than raw throughput.
#Go further
The DGX Spark is best appreciated by comparing it with the alternatives already covered on the site. These guides extend the discussion:
- Which LLM on RTX 5090 (32 GB)?
- A detailed comparison of the fastest consumer card, to weigh throughput against capacity.
- Which LLM on Mac Studio?
- The other major unified-memory machine, up to 512 GB—the true rival to the DGX Spark.
- Choose your GPU for local AI
- To scope your VRAM needs before spending money and see where each option fits.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.