Family Nemotron · 30B parameters

Nemotron TwoTower 30B-A3B Base

30B/3B-active MoE base model (NVIDIA). 128k context, ~17 GB VRAM in Q4. Foundation for fine-tuning and reasoning.

🇺🇸 NVIDIA·License NVIDIA Open Model License·Context 125k tokens·Output 2026-04-11·Fits within the 128 GB of the GIGABYTE AI TOP ATOM← Catalog

01What it can do

Strengths
  • MoE 30B/3B active: lightweight inference for its size
  • 128k native context
  • Ideal foundation for custom fine-tuning
  • Commercial use permitted (NVIDIA Open Model License)
Limitations to know
  • —Unaligned base model—requires fine-tuning/instruct for chat
  • —No official Ollama tag — install via HuggingFace
  • —~17 GB VRAM minimum in Q4
Architecture
30B total MoE / 3B active · BF16 weights · TwoTower architecture from Nemotron Labs
Training
Base (non-instruct) model from the Nemotron Labs family. Native 128k-token context.
Ideal for
Fine-tuningReasoningCode

04Install

Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.

$# HuggingFace : nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16
⚠
First download: between 2 and 40 GB depending on the selected quantization. Plan for sufficient disk space; a stable connection is recommended. Subsequent launches are instant.

02Required memory

Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.

Q4_K_M
The lightest, ~5% loss
17 GB
Q5_K_M
Good quality/size compromise
21 GB
Q8_0
Nearly indistinguishable from FP16
32 GB
FP16
Full precision — server use
60 GB
Fallback CPU · If you don't have a GPU, allow 39 GB of RAM minimum to run this model at reduced speed.

What hardware do you need for Nemotron TwoTower 30B-A3B Base?

To run Nemotron TwoTower 30B-A3B Base locally with Q4 quantization, you need about 17 GB of VRAM. An option to compare: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — leave some headroom for the system and context; check engine compatibility with the GPU.

Current offer: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395)
AmazonSee price →

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — possible commission at no extra cost to you; independent recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

On the go: Nemotron TwoTower 30B-A3B Base also runs on a RTX laptop PC (24 GB of VRAM) →

03Expected speed

Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.

Entry-level
~9t/s
GTX 1650, RX 6600, MBA M2 8GB
Mid-range
~14t/s
RTX 4060, 4070, MBP M3 Pro
High-end
~22t/s
RTX 4090, M4 Max, Radeon 7900