Family Nemotron · 561B parameters

Nemotron 3 Ultra (BF16)

⚠ Nemotron 3 Ultra BF16 (561B MoE, ~55B active). 128k context, multilingual reasoning and coding. Datacenter (~325 GB Q4 VRAM). NVIDIA Open Model License.

🇺🇸 NVIDIA·License NVIDIA Open Model License·Context 125k tokens·Output 2026-06-03← Catalog

01What it can do

Strengths
  • MoE frontier: 561B / 55B active in BF16
  • Multilingual, 11 languages (including French)
  • 128k native context
  • NVIDIA Open Model License (commercial use OK)
Limitations to know
  • —~325 GB VRAM in Q4, ~1.1 TB in BF16 — H100/MI300 multi-GPU mandatory
  • —No Ollama tag: install via Hugging Face only
  • —NVIDIA Open Model License (review based on commercial use)
Architecture
Mixture-of-Experts · 561B total parameters · ~55B active per token · 128k context · BF16 weights
Training
BF16 checkpoint (post-Base variant) from the Nemotron 3 Ultra family. Multilingual pretraining covering 11 languages (en, fr, es, it, de, pt, ja, ko, hi, ar, zh).
Ideal for
Datacenter reasoningMultilingual codeHigh-capacity MoE

04Install

Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.

$# HuggingFace : nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
⚠
First download: between 2 and 40 GB depending on the selected quantization. Plan for sufficient disk space; a stable connection is recommended. Subsequent launches are instant.

02Required memory

Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.

Q4_K_M
The lightest, ~5% loss
325 GB
Q5_K_M
Good quality/size compromise
398 GB
Q8_0
Nearly indistinguishable from FP16
600 GB
FP16
Full precision — server use
1122 GB
Fallback CPU · If you don't have a GPU, allow 729 GB of RAM minimum to run this model at reduced speed.

What hardware do you need for Nemotron 3 Ultra (BF16)?

To run Nemotron 3 Ultra (BF16) locally with Q4 quantization, you need about 325 GB of VRAM. An option to compare: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) — this model exceeds this mini-PC's GPU capacity: choose a smaller model or suitable infrastructure.

Current offer: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395)
AmazonSee price →

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Affiliate links — commission possible at no extra cost to you; independent recommendation. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

03Expected speed

Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.

Entry-level
~1.5t/s
GTX 1650, RX 6600, MBA M2 8GB
Mid-range
~2.5t/s
RTX 4060, 4070, MBP M3 Pro
High-end
~5t/s
RTX 4090, M4 Max, Radeon 7900