gpt-oss locally: OpenAI’s open-weight models with Ollama
OpenAI eventually released open weights: gpt-oss-20b and gpt-oss-120b, two mixture-of-experts models under the Apache 2.0 license. This guide shows how to run gpt-oss with Ollama, how much VRAM each version requires, how native MXFP4 quantization works, and how to tune the reasoning effort. We conclude by looking at how these open weights really compare with Qwen and DeepSeek.
#Why gpt-oss changes the game
For years, OpenAI kept its models closed. gpt-oss is the company's first return to genuinely downloadable weights since GPT-2, this time under the Apache 2.0 license: commercial use allowed, no forced-sharing clause, and no user limit. You download the weights, they run on your machine, and nothing leaves your network.
Technically, gpt-oss uses a MoE (Mixture of Experts) architecture: the model contains many total parameters but activates only a small fraction for each token. The result is quality close to that of a large dense model, with the memory cost and speed of a much smaller model. Both versions target reasoning, tool calling, and instruction following, with a context of 128k tokens.
- gpt-oss-20b
- ≈ 21 billion total parameters, ≈ 3.6B active per token. Designed for a single 16 GB consumer card.
- gpt-oss-120b
- ≈ 117 billion parameters in total, ≈ 5.1B active per token. Intended for an 80 GB data-center card or a system with a large amount of unified memory.
- License
- Apache 2.0 — commercial, redistributable, and fine-tunable without usage restrictions.
- Quantization
- Native MXFP4 on expert weights: the 4-bit format is part of the release, not an approximate third-party conversion.
#gpt-oss ollama: 20b or 120b?
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The choice depends almost entirely on your hardware. Both models belong to the same family and use the same response formatting; the quality difference exists but is secondary to the question, “Will it fit in my VRAM?” Here’s the simple rule of thumb.
- 20B—the reasonable default
- About 14–16 GB once loaded in MXFP4. Fits on a RTX 4080 16 GB or a RTX 4090, and on a Apple Silicon Mac starting at 24 GB of unified memory. This is the version to install first.
- 120B — when you have the memory
- About 60–65 GB in MXFP4. It targets an 80 GB GPU (H100/A100 class), a multi-GPU workstation, or a Mac with a large amount of unified memory (M3/M5 Ultra 96 GB and up). There is no point trying it with less than 64 GB of memory available to the model.
- Relative quality
- The 120b widens the gap in multi-step reasoning, long-form code, and ambiguous tasks. For everyday question answering and French, the 20b already performs very well.
#Prerequisites
Nothing exotic: an up-to-date Ollama installation and enough memory for the target version. Ollama handles the download, MXFP4 quantization, and GPU/CPU offloading automatically.
- Ollama recent
- Version supporting gpt-oss and MXFP4. Update before getting started—older versions don’t recognize the format.
- VRAM / unified memory
- ≥ 16 GB for the 20b, ≥ 64 GB for the 120b. Below that, Ollama will spill into CPU RAM and speed will collapse.
- Disk space
- ≈ 12-14 GB for the 20b, ≈ 60-65 GB for the 120b. The weights are stored in the Ollama model directory.
- Daemon listening
- Ollama exposes its API on http://localhost:11434 by default—convenient for connecting Open WebUI or building code on top of it.
#Install gpt-oss-20b with one command
The 20b version installs and launches in a single line. Ollama downloads the MXFP4 weights, loads them onto the GPU, and opens a chat session.
On the first launch, allow time to download ~13 GB. The model then remains cached. To just retrieve the weights without opening an interactive session, use pull.
In ollama ps, you want to see a high GPU percentage. If the column shows « 100% CPU », the model did not fit in VRAM and spilled over — generation will be slow. On a RTX 4090, expect comfortable interactive throughput; on a 4080 16 GB, the 20B fits but leaves no room for a large context.
To control the model through an API rather than chat, the endpoint is already available. Here is a minimal call.
#Which machines for the 120b
The command is identical; only the tag changes. But run it only if you actually have enough memory: below 64 GB, Ollama will offload some experts to the CPU, and you will get a few tokens per second at best.
- 01Check available memoryOn a NVIDIA GPU, nvidia-smi should show an 80 GB card (or multiple cards with enough combined VRAM). On a Mac, 96 GB+ of unified memory leaves room for the model and the system.
- 02Start the pullThe download is substantial (~60 GB): plan for the bandwidth and disk space. The MXFP4 format eliminates the need to requantize it yourself.
- 03Control placementAfter ollama run, an ollama ps in another terminal confirms the GPU usage. Massive CPU offloading = insufficient memory; go back to 20b.
- 04Adjust the contextThe 120b supports 128k tokens, but a full context consumes a lot of additional memory. Reduce num_ctx if you are short on VRAM.
#Adjust reasoning effort
What’s distinctive about gpt-oss: the models produce a chain of reasoning before their answer, and you can control how much effort they spend. Three levels—low, medium, high—balance speed and depth. Low effort answers quickly for simple tasks; high effort thinks longer for math, code, or multi-step logic.
The setting is configured in the system message. Specify the desired level, and the model adjusts the length of its internal reasoning.
- Reasoning: low
- Fast responses, little visible reasoning. Ideal for rephrasing, short summaries, and factual questions.
- Reasoning: medium
- Good default compromise: some reasoning without latency exploding.
- Reasoning: high
- In-depth reasoning, better on hard problems—at the cost of longer generation and more tokens consumed.
#Strengths and weaknesses compared with Qwen and DeepSeek
gpt-oss enters a field already occupied by very capable open-weight models. Here's where it stands out and where it falls short, without sugarcoating.
- Strengths
- Careful reasoning and tool calling, adjustable effort, a straightforward Apache 2.0 license, and a controlled memory footprint thanks to MoE + MXFP4. The 20b delivers an outstanding quality-to-VRAM ratio on a single 16 GB card.
- Compared with Qwen3
- Qwen often remains ahead in multilingual tasks and French, with a broader range of sizes (from 4B to large MoE models). gpt-oss makes up for it with the quality of its reasoning chain and clean instruction following.
- Compared with DeepSeek
- The DeepSeek reasoning models (R1 and V3.2) go further on very difficult problems, but their larger versions require much more hardware. gpt-oss 20b targets accessible local use, while DeepSeek V3.2 targets the high end.
- Weakness to know about
- French is not gpt-oss’s priority: quality is good but slightly below English, and sometimes below Qwen. Test it on your own prompts before deciding.
#Troubleshooting
- “unknown model” or unrecognized format
- Your Ollama is too old for MXFP4. Update to a recent version, then run the pull again.
- Very slow generation on the 20b
- ollama ps montre du CPU au lieu du GPU : le modèle a débordé. Fermez les autres applis gourmandes en VRAM, réduisez num_ctx, ou vérifiez que le GPU est bien détecté.
- The 120b crashes or crawls painfully
- Insufficient memory. With less than 64 GB actually available for the model, stick with the 20b—the 120b is not designed for your machine.
- Responses are too slow for everyday use
- Lower the reasoning effort to low or medium in the system message: high uses an enormous number of reasoning tokens.
- Context that blows past VRAM
- 128k tokens are expensive in memory. Set num_ctx to a reasonable value (for example, 8192) for your everyday use.
#Go further
Once gpt-oss is in place, these site guides help you fine-tune your local stack and compare models:
- Choose your quantization (Q4, Q5, Q8, FP16)
- To understand where MXFP4 fits relative to classic GGUF formats and make quality/memory trade-offs for your other models.
- Choose your GPU for local AI
- The 2026 buying guide for sizing the card depending on whether you're targeting 20b on 16 GB or more demanding models.
- Install DeepSeek R1 with Ollama
- The direct, real-world comparison of reasoning: install R1 and evaluate it side by side with gpt-oss.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.