01What it can do
- MoE 30B/3B active: lightweight inference for its size
- 128k native context
- Ideal foundation for custom fine-tuning
- Commercial use permitted (NVIDIA Open Model License)
- —Unaligned base model—requires fine-tuning/instruct for chat
- —No official Ollama tag — install via HuggingFace
- —~17 GB VRAM minimum in Q4
04Install
Install Ollama for your OS. Check the model and its quantization before downloading. Start with 4096 tokens of context, then check placement with ollama ps. A command below is not proof that a test was run on your machine.
02Required memory
Approximate GPU VRAM required to run this model, including 4k tokens of context overhead. For a longer context, add ~1 GB per 8k-token increment.
What hardware do you need for Nemotron TwoTower 30B-A3B Base?
To run Nemotron TwoTower 30B-A3B Base locally with Q4 quantization, you need about 17 GB of VRAM. An option to compare: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — leave some headroom for the system and context; check engine compatibility with the GPU.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Affiliate links — possible commission at no extra cost to you; independent recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
On the go: Nemotron TwoTower 30B-A3B Base also runs on a RTX laptop PC (24 GB of VRAM) →
03Expected speed
Tokens generated per second in Q4_K_M, 4k context. Beyond 20 t/s, reading is comfortable. Below 10 t/s, that's just for testing.