Intermediate 14 minMini-PC

NVIDIA DGX Spark: the 128 GB AI mini-PC versus GPU

The DGX Spark is NVIDIA's bet on local desktop AI: a Grace Blackwell chip, 128 GB of unified memory, and a form factor that fits on a corner of your desk. On paper, it runs models that no consumer GPU can load on its own. This guide takes a practical look at the dgx spark llm: what fits in memory, real inference speeds, and a head-to-head comparison with a RTX 5090 and a Mac Studio to determine whether this mini PC deserves a place in your setup.

Choosing a machine? Our picks by budget → · Our spec sheet NVIDIA DGX Spark →

By Samir K.·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What the DGX Spark is

The DGX Spark is a compact computer built around the NVIDIA GB10 Grace Blackwell superchip. It combines a Blackwell GPU and a Grace ARM CPU (20 cores) on the same chip, connected to a single pool of 128 GB of LPDDR5X memory shared between the two. This unified-memory architecture—the same philosophy as Apple Silicon—is what makes the machine so compelling for local AI.

Where a conventional graphics card isolates its VRAM (24 or 32 GB maximum for consumer hardware), the DGX Spark lets the GPU address all 128 GB directly. In practice, the constraint that blocks everyone on GPUs—“does the model fit in VRAM?”—almost disappears. The limiting factor becomes memory bandwidth, not capacity.

Chip
NVIDIA GB10 Grace Blackwell (Blackwell GPU + 20-core ARM CPU)
Memory
128 GB unified LPDDR5X, shared by the CPU/GPU
Bandwidth
~273 GB/s (LPDDR5X), the real bottleneck
Network
ConnectX (200 GbE) to link two Sparks and target very large models
Format
Desktop mini PC, ~150 W, quiet compared with a multi-GPU tower
i
This is not a datacenter replacement
The DGX Spark is aimed at prototyping, light fine-tuning, and local inference of large models—not large-scale training or high-load serving. NVIDIA positions it as a personal AI development workstation.

#128 GB unified memory: what it unlocks

For context, here is the approximate VRAM required by a model in Q4_K_M quantization, the most common choice for local inference: a 7B fits in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B in ~40 GB. On a RTX 5090 with its 32 GB, a 70B Q4 therefore does not fit entirely: part of it must be offloaded to the CPU, which causes a dramatic drop in speed.

With 128 GB of unified memory, the DGX Spark easily loads a 70B model in Q4, Q5, or even Q8, leaves room for a long context (32k tokens and more), and opens the door to 100B+ parameter models with aggressive quantization. This is the class of model that, until now, required either a high-end Mac Studio or an expensive, power-hungry multi-GPU rig.

70B Q4 (~40 GB)
Comfortable, with plenty of room for context
70B Q8 (~75 GB)
Possible—near-FP16 quality on a large model
120B-type MoE in Q4
Fits in memory; bandwidth determines speed
Multiple models loaded
One 32B + one 14B in parallel for routing or multi-agent workloads
→
Memory isn’t speed
Loading a model and running it quickly are two different things. The 128 GB ensures the model fits; the roughly 273 GB/s of bandwidth determines tokens/sec. With very large models, expect smooth inference for reading, not bursts of generation.

#Which models it runs

The DGX Spark's strength is the 30B–120B range, exactly where consumer GPUs stall due to insufficient VRAM. Here are the typical uses by model family.

Qwen 3.8 27B
The flagship generalist of 2026 (vision, 262k ctx); on this machine it runs at full-quality Q8 with plenty of context headroom
Qwen 3.6 35B-A3B / Qwen3-Coder 30B
Fast MoEs for reasoning and coding, fully resident in memory even in Q8
GLM 4.7 Flash (MoE 30B-A3B, MIT)
Excellent for agents, loaded at full quality without juggling
Large 100B+ quantized MoE
Mixture-of-Experts architectures fully benefit from the large unified memory
Multimodal models
Gemma 4 and Qwen 3.8 (vision + text), memory-hungry, run without compromise

For small models (7B–14B), the DGX Spark obviously works, but it's overkill: a RTX 4070 or an M4 Mac mini runs them faster for far less money. The machine only makes sense if you regularly target large models.

i
GB10: a datacenter superchip at home
The Spark's core is the GB10 Grace Blackwell: 20 ARM cores (10 Cortex-X925 + 10 A725) soldered to a Blackwell GPU through an NVLink-C2C link, ~1 petaflop in FP4, and 128 GB of unified memory. It is literally the architecture of GB200 servers reduced to paperback size—with the full CUDA ecosystem, making it the only direct bridge between “experimenting at home” and “deploying in production.”

#Real-world inference performance

Two figures matter in local inference: generation speed (tokens/sec, what scrolls across the screen) and prefill time (the delay before the first token, which depends on prompt length). Memory bandwidth primarily drives generation.

On the DGX Spark, a 70B model in Q4 typically generates around ten tokens/sec—enough for a comfortable assistant experience where you read as output arrives, but less suitable for high-throughput batch workloads. Smaller models (14B–32B) go significantly faster. This profile is consistent with ~273 GB/s of bandwidth: that's the wall separating the DGX Spark from much faster GDDR7 cards in this respect.

!
Don't compare tokens/sec out of context
A tokens/sec figure is meaningful only with the exact model, quantization, context size, and inference engine. The values above are rough estimates for orientation, not guarantees. Measure your own workload before making a purchase decision.

#DGX Spark vs RTX 5090

This is the most revealing comparison, and it pits two philosophies against each other. The RTX 5090 offers 32 GB of GDDR7 VRAM with ~1792 GB/s of bandwidth—about six to seven times faster than the DGX Spark's memory. For anything that fits within its 32 GB, the 5090 crushes the Spark in tokens/sec.

But the 5090 hits a hard wall: beyond ~32 GB, it has to offload to the CPU and performance collapses. A 70B Q4 (~40 GB) doesn’t fit entirely in it. The DGX Spark, by contrast, handles that same 70B without flinching—more slowly, but as a single piece. The choice comes down to one question: does your workload fit within 32 GB?

You stay at ≤ 32 GB (up to ~32B Q4)
RTX 5090: much faster and versatile (gaming, rendering, CUDA)
You are targeting 70B and above
DGX Spark: the model fits where the 5090 gives up
Power consumption / noise
DGX Spark: ~150 W and quiet, versus a hot, noisy 5090 tower
Complete CUDA ecosystem
Both; the Spark adds the DGX/NVIDIA stack end to end

#DGX Spark vs. Mac Studio

The DGX Spark's real rival isn't a graphics card; it's the Mac Studio. Both share the same idea—massive unified memory in a compact, efficient enclosure—and target the same audience: developers who want large models locally without building a rig.

The Mac Studio (M-Ultra) supports up to 512 GB of unified memory with often higher bandwidth on high-end configurations, making it formidable for very large models. Its advantage: macOS, Metal, and a mature ecosystem (Ollama, LM Studio, MLX). The DGX Spark's advantage: native CUDA. If your pipeline depends on CUDA libraries, TensorRT, or NVIDIA tools, the Spark natively speaks the same language as your production servers.

Maximum memory capacity
Mac Studio up to 512 GB; DGX Spark 128 GB (expandable by chaining 2 units)
Software ecosystem
Mac: mature Metal/MLX; Spark: native CUDA, continuity with NVIDIA servers
AI code portability
The Spark runs code written for NVIDIA cloud/datacenter GPUs without friction
Office use
The Mac Studio is a complete workstation; the Spark is a dedicated AI station
i
The criterion that often decides
If you prototype code intended to run on NVIDIA GPUs in production, the DGX Spark eliminates porting surprises. If you mainly want comfortable local inference and a versatile machine, the Mac Studio is more flexible. The rest comes down to bandwidth and budget.

#Install Ollama on it

The DGX Spark runs on a Linux base (DGX OS, derived from Ubuntu) with the CUDA stack preinstalled. Ollama works there like on any Linux machine NVIDIA: the daemon detects the GPU and serves models on its usual port.

  1. 01
    Install Ollama
    Official one-command script. It configures the systemd service and automatically detects the Blackwell GPU through CUDA.
  2. 02
    Check GPU detection
    Confirm that the GPU is detected before launching a large model; otherwise, inference would fall back to the CPU.
  3. 03
    Pull a large model
    Download a large 2026 MoE at full quality (a Qwen3-Coder 30B in Q8, or the flagship Qwen 3.8 27B with its 262k context) to use the available memory—this is exactly what the machine is designed for.
  4. 04
    Launch and connect an interface
    The daemon listens on http://localhost:11434 by default; point Open WebUI or LM Studio to it.
Terminal — installation
# Installer Ollama (Linux / DGX OS)
curl -fsSL https://ollama.com/install.sh | sh

# Vérifier que le GPU Blackwell est détecté
nvidia-smi

# Tirer et lancer un gros MoE 2026 en pleine qualité (exploite la mémoire unifiée)
ollama run qwen3-coder:30b-a3b-q8_0

The Ollama daemon exposes a compatible API on port 11434. To make the most of your memory, favor models that standard GPUs cannot load on their own, and increase the context window since you have room for it. Note that Qwen 3.8 27B tends to overthink with its default reasoning setting: switch it to low mode for more direct answers.

Terminal — extended context
# Servir Qwen 3.8 27B avec un grand contexte depuis un Modelfile
cat > Modelfile <<'EOF'
FROM qwen3.8:27b
PARAMETER num_ctx 131072
EOF

ollama create qwen38-128k -f Modelfile
ollama run qwen38-128k
→
Chain two Sparks
The 200 GbE ConnectX network port lets you link two DGX Spark systems to combine their memory and target models above 128 GB. This is the upgrade path planned by NVIDIA if a single box is no longer enough.

#One chip, eight machines: certified clones

NVIDIA turned GB10 into a platform: besides its DGX Spark Founders Edition, seven manufacturers sell the same certified base under their own brands. Same chip, same 128 GB—the differences come down to storage, chassis, price, and sales channel.

The 8 certified GB10 systems (August 2026)
MachineManufacturerKey takeaway
DGX Spark Founders EditionNVIDIAThe benchmark, ~$3,999
Ascent GX10AsusThe best price point (~$2,999)
Veriton GN100AcerSingle SKU with 4 TB of storage
AI TOP ATOMGigabyteComes with the AI TOP software suite
EdgeXpert MS-C931MSIIndustrial/edge channel
Pro Max GB10DellEnterprise channel
ThinkStation PGXLenovoWorkstation channel
ZGX NanoHPWorkstation channel

A detail that changes everything for large models: the integrated ConnectX port lets you link two units and run 400B-class models locally. No other consumer machine offers this.

#Frequently asked questions

What exactly is the DGX Spark's GB10?+
A “superchip”: a 20-core ARM Grace processor and a Blackwell GPU fabricated to work together, connected by NVLink-C2C (much faster than a PCIe bus), sharing 128 GB of unified LPDDR5X memory. It delivers about 1 petaflop in FP4—the architecture of NVIDIA’s AI servers at mini-PC scale.
What is the difference between the DGX Spark and its Asus, Acer, or Gigabyte clones?+
None on the chip: all 8 systems use the same GB10 and the same 128 GB. The differences are storage (up to 4 TB on Acer), the chassis, bundled software, support, and especially the price—as the Asus Ascent GX10 is about $1,000 cheaper than the Founders Edition.
Can you really run a 400B model on two machines?+
Yes, this is an official use case: two units connected through their ConnectX 200 Gb/s ports pool their memory to serve 400B-class models (very large 400B+ class MoE models in quantized form). This is unique in the consumer world—the next step is a server costing several tens of thousands of euros.
DGX Spark or Strix Halo machine?+
Both offer 128 GB of unified memory and comparable text generation performance (their memory bandwidths are close). The GB10 wins on prompt processing, fine-tuning, and the CUDA ecosystem; Strix Halo wins on price (often half as expensive), Windows, and versatility. Our dedicated comparison decides profile by profile — verdict: tie.
Who is a GB10 machine for?+
For AI developers who want to prototype locally what they will deploy on NVIDIA in production, teams doing fine-tuning, and those targeting very large models through clustering. For everyday chat and coding assistance, a Strix Halo machine delivers similar performance for half the price.

#Who this format makes sense for

The DGX Spark is not a consumer machine. Its value is concentrated among specific users for whom the VRAM constraint is the real day-to-day problem.

AI developers on CUDA
Locally prototype code intended for NVIDIA production GPUs, without porting surprises
Researchers and R&D teams
Light fine-tuning and inference of large models without reserving cloud GPU time
70B+ power users
Those who constantly run into the 24–32 GB limit of a consumer graphics card
Noise- and power-sensitive offices
A quiet, efficient workstation rather than a multi-GPU rig

Conversely, if you game, mainly run models ≤ 32B, or have a tight budget, a RTX 5090 (or even a 4070/4080) or an M4 Mac mini will offer much better performance per dollar. The DGX Spark makes sense when massive memory capacity matters more than raw throughput.

!
The magic-number trap
“128 GB” sounds impressive, but do not rent them to run an 8B. Buy this machine for what it truly unlocks — the 70B–120B range locally — or you will pay dearly for capacity you never use.

#Go further

The DGX Spark is best appreciated by comparing it with the alternatives already covered on the site. These guides extend the discussion:

Which LLM on RTX 5090 (32 GB)?
A detailed comparison of the fastest consumer card, to weigh throughput against capacity.
Which LLM on Mac Studio?
The other major unified-memory machine, up to 512 GB—the true rival to the DGX Spark.
Choose your GPU for local AI
To scope your VRAM needs before spending money and see where each option fits.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.