Intermediate 10 minNVIDIA

Nemotron 3 locally: NVIDIA models with Ollama

NVIDIA does not only make GPUs: the brand also post-trains its own LLMs, the Nemotron family. This guide shows how to install Nemotron 3 locally with Ollama, which variants (Nano, Super) to choose based on your VRAM, what NVIDIA optimizes differently from the others—controllable reasoning and agentic use—and when these models are worth using over equivalent Qwen3 models.

By Clara M.·Update 2026-07-28·Tested on Windows, macOS, and Linux

#Why Nemotron instead of another model

Nemotron isn't a model trained from scratch. NVIDIA starts from solid open-weight foundations (Llama, Qwen depending on the version), then applies its own post-training recipe: pruning to reduce size, distillation to recover lost quality, and massive fine-tuning focused on reasoning, mathematics, code, and tool calls. The result targets a specific point: the best quality-to-compute-cost ratio for agent tasks, not general-purpose conversation.

The other distinctive feature is that NVIDIA optimizes these models for its own hardware and inference pipelines. In practice, this means solid throughput on NVIDIA GPUs and a family designed as a coherent range: a small version for consumer cards, a medium version for workstations, and a giant version for servers. You choose based on available VRAM without changing your prompting approach.

i
Nemotron is improved Llama/Qwen
If you already know Llama or Qwen locally, Nemotron behaves familiarly: the same chat format and the same GGUF quantizations. The difference lies in the NVIDIA post-training, not in an exotic architecture.

#Nano, Super, and Ultra: which variant

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The Nemotron 3 family comes in three sizes, each intended for a category of machine. The name indicates the target more than the exact parameter count.

Nano (~8–9B)
The version for consumer GPUs. Easily fits on a RTX 3060 12 GB or any GPU with at least 8 GB of VRAM in Q4. It is the ideal entry point for testing Nemotron and lightweight agents.
Great (~49B)
The version for workstations and large GPUs. It targets the quality of a 70B model at a lower compute cost, thanks to NVIDIA pruning. It requires a RTX 4090 24 GB (with tight Q4) or, preferably, a 32–48 GB card.
Ultra (~253B)
The server version, beyond the reach of a consumer machine. Reserved for multi-GPU workstations or the data center. Mentioned here to put the range in context, not for typical local use.

For a local workstation, the real choice is between Nano and Super. Choose Nano if you have an 8–16 GB card or if speed comes first. Choose Super if you have 24 GB or more and want maximum quality for demanding reasoning or coding.

#Requirements and VRAM by variant

As with any local model, VRAM is the limiting factor. Here are the benchmarks for Q4_K_M quantization, the default recommendation because it offers the best quality/memory tradeoff. Always leave some headroom for the context window, which consumes additional memory beyond the weights.

Nemotron Nano (~9B) — Q4
≈ 6-7 GB of VRAM. Comfortable on RTX 3060 12 GB, RTX 4070 12 GB, or a Mac Apple Silicon with at least 16 GB of unified memory.
Nemotron Super (~49B) — Q4
≈ 28-30 GB. It does not fit on a single RTX 4090 24 GB without offload: plan for a 32 GB+ card, two GPUs, or an M4 Pro/Max Mac with 48 GB of unified memory.
Nemotron Super — Q5/Q8
Q5_K_M climbs to around 35 GB, while Q8_0 exceeds 50 GB. Reserved for high-VRAM or multi-GPU configurations.
Context headroom
Add 1 to 4 GB depending on the target context length. A long reasoning context generates many intermediate tokens, so it increases the memory footprint.
→
Verify what actually fits
After a “ollama run”, run “ollama ps” in another terminal. The PROCESSOR column shows “100% GPU” if everything is in VRAM, or a “GPU/CPU” split if Ollama had to spill into RAM—in which case speed drops sharply. That's the signal to step down in size or quantization level.

On the software side, the prerequisite is minimal: the Ollama daemon installed and listening by default on http://localhost:11434. On NVIDIA GPUs, make sure the CUDA drivers are up to date; Ollama detects the GPU automatically. An interface such as Open WebUI or LM Studio on top is convenient, but not required.

#Install Nemotron via Ollama

Installation follows the usual Ollama workflow exactly. If the daemon is already running, a single command downloads and starts the model.

  1. 01
    Verify that Ollama is running
    Open a terminal and run “ollama --version”. If the command responds, the daemon is ready. Otherwise, install Ollama first (see the installation guide at the end of the article).
  2. 02
    Download the Nano variant
    To test without risking your VRAM, start with Nano. The command downloads the weights and then opens a chat session directly in the terminal.
  3. 03
    Upgrade to Super if your GPU can handle it
    Once you're comfortable, and if you have enough VRAM, pull the Super variant for better reasoning and coding quality.
  4. 04
    List and manage models
    “ollama list” displays the installed models and their disk size. “ollama rm <model>” frees the space used by a variant you no longer use.
Terminal — Nemotron installation
# Vérifier qu'Ollama répond
ollama --version

# Variante Nano (~9B) — cartes grand public
ollama run nemotron-nano

# Variante Super (~49B) — grosse VRAM / multi-GPU
ollama run nemotron-super

# Voir ce qui est installé et vérifier le placement GPU
ollama list
ollama ps
!
The tags evolve
The Ollama library regularly updates the names and versions of Nemotron models. Before running a command, check the exact tag on ollama.com/library—the model name and version number may differ from the example above depending on the published generation.

To call Nemotron from your own scripts, Ollama's REST API is available on the same port. The “model” field uses the exact tag you pulled.

Terminal — API call
curl http://localhost:11434/api/generate -d '{
  "model": "nemotron-nano",
  "prompt": "Explique la différence entre RAG et fine-tuning en trois phrases.",
  "stream": false
}'

#Controllable reasoning mode

Nemotron models are characterized by reasoning that you can enable or disable on demand. NVIDIA trained these models to produce a detailed chain of thought when asked via the system prompt, and to answer directly otherwise. In practice, you switch between two modes with a simple system instruction: “detailed thinking on” to force step-by-step reasoning, “detailed thinking off” for a concise, fast response.

This switch is useful because reasoning has a cost. A chain of thought generates many intermediate tokens: that’s valuable for a math or logic problem or subtle debugging, but wasteful for rewriting an email. With Nemotron, you decide prompt by prompt whether to pay that overhead.

Terminal — enable reasoning
# Dans une session ollama run, définir le system prompt :
/set system detailed thinking on

# Puis poser une question de raisonnement
Un train part de A à 60 km/h, un autre de B à 90 km/h... résous étape par étape.
→
Turn off reasoning to go faster
For simple text, rewriting, or classification, turn “detailed thinking off” on. You gain latency and tokens with no loss of quality on tasks where thinking out loud adds nothing.

This brings Nemotron closer to models such as DeepSeek R1 or Qwen3 in their “thinking” mode, but with more explicit control: reasoning is not a global model mode; it is a setting you control throughout the conversation.

#Code, agents, and tool calling

NVIDIA primarily optimizes Nemotron for two use cases: reasoning and agentic workflows. On code, the Super variants hold their own against good 32B–70B models for generation, explanation, and correction. Nano remains adequate for occasional assistance but shows its limits on long tasks or major refactors—that is expected for a model of its size.

The real differentiator is function calling. Nemotron is trained to produce reliable, well-formed tool calls, making it a good candidate for building local agents: a model that decides to call a function, reads the result, and continues. Combined with reasoning mode, this enables agents to plan before acting instead of rushing ahead.

Reasoning / math
A clear strength, especially on Super with “detailed thinking on.” The chain of thought reduces calculation and logic errors.
Code
Super rivals 70B models for generation and debugging; Nano is suitable for lightweight assistance and autocompletion.
Agents and tools
Reliable tool calls, compliant JSON format—one of the family’s best uses for local automation.
French / general-purpose conversation
Good without being the best in its category. Nemotron is built for reasoning and acting, less for multilingual chatter.

#Against Qwen3: when to prefer Nemotron

Qwen3 is the open-weight benchmark to compare against, because both families operate in the same space: reasoning, code, multiple sizes, and thinking mode. On many general-purpose tasks and in French quality, Qwen3 remains ahead at the same size—it is a more versatile model with better multilingual training.

Nemotron pulls ahead when your use case leans toward agentic workflows and integration with a NVIDIA ecosystem. If you are building an agent with many tool calls, want explicit control over reasoning prompt by prompt, or use a NVIDIA GPU cluster where inference throughput matters, Nemotron is worth testing. Conversely, for a general-purpose conversational assistant in French, Qwen3 is the safer default choice.

Choose Nemotron if
Agents with intensive tool calling, a need to finely switch reasoning on or off, 100% NVIDIA infrastructure, or a search for the best quality-to-compute ratio for pure reasoning.
Choose Qwen3 if
General-purpose assistant, with French and multilingual use as priorities, conversational versatility, or simply the most proven model for broad use.
Practical verdict
Test both at comparable sizes (Nano vs Qwen3-8B, Super vs Qwen3-32B) on YOUR real prompts. The gap depends heavily on the task, not on an abstract ranking.

#Troubleshooting

“model not found” on pull
The tag does not exist under this name. Check the exact spelling on ollama.com/library—the Nemotron names change from one generation to the next.
Speed collapses on Super
The model spills into RAM because of insufficient VRAM. “ollama ps” shows GPU/CPU splitting. Switch to Nano, use lower quantization (Q4), or add a GPU.
Reasoning does not activate
The system prompt is not being applied. In an interactive session, use “/set system detailed thinking on”; in the API, pass the instruction in the system field, not in the user prompt.
Overly verbose responses
Reasoning mode is enabled by default or inherited. Switch to “detailed thinking off” for simple tasks.
“Connection refused”
The Ollama daemon is not running. Check with “ollama ps” and make sure the service is listening on http://localhost:11434.

#Go further

Nemotron builds on components already covered on the site. These guides extend this one:

Qwen 3 locally: complete testing and real-world benchmarks
The essential comparison for positioning Nemotron against its main open-weight alternative, with tokens-per-second figures.
Choose your quantization (Q4, Q5, Q8, FP16)
To balance quality and VRAM on Nemotron Super, where each quantization level changes what fits on your card.
Install Ollama on Linux
The prerequisite if the daemon isn't set up yet: install script, systemd service, and NVIDIA GPU configuration.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.