Install Ollama on Windows 11: complete guide (2026)
On Windows 11, Ollama installs through a standard installer downloaded from the official website, with automatic GPU detection (native CUDA NVIDIA support). A single terminal command (ollama run qwen3.5:4b) downloads and launches a first model in about 5 minutes, with no manual configuration or complex command line.
You have a Windows PC and want to run an LLM locally without using ChatGPT. Ollama is probably the simplest tool for the job: a standard installer, one command, one model. In exactly 3 minutes, you’ll have Qwen 3.5 running on your machine.
#Why Ollama?
Ollama handles all the tedious work for you: downloading models, quantization, loading into VRAM, GPU detection, and the local HTTP API. Under the hood, it’s llama.cpp. But you won’t have to compile it.
The tool is free, open source, and runs entirely offline once the model has been downloaded. No data is sent to a remote server.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Windows 10 or 11
- 64-bit. Ollama has no 32-bit version.
- 8 GB of RAM minimum
- 16 GB is comfortable for 8B to 9B models in Q4 quantization.
- 10 GB of disk space
- A 4B to 9B model weighs about 3 to 7 GB. Leave plenty of room if you plan to test several.
- GPU NVIDIA (optional)
- Up-to-date drivers, RTX 20/30/40/50 or GTX 16. Ollama detects the GPU automatically.
#1. Download
Go to the official website and download the Windows installer.
The installer is about 700 MB. That's normal: it includes precompiled llama.cpp binaries with CUDA support.
#2. Installation
- 01Run OllamaSetup.exeDouble-click the downloaded installer. Windows Defender may display a SmartScreen warning — click "More info," then "Run anyway."
- 02Silent installationNo configuration window: the installer installs to %LOCALAPPDATA%\Programs\Ollama and automatically starts the service in the background.
- 03Check the system tray iconA small Ollama icon appears near the clock in the bottom-right corner. It indicates that the daemon is running.
- 04Open a terminalPowerShell, Windows Terminal, or cmd.exe—it doesn't matter. Type the verification command below.
You should see something like ollama version is 0.5.x. If the command isn't recognized, close and reopen the terminal—the PATH was just modified.
#3. Your first model
Qwen 3.5 4B is a good first choice: multilingual (including very good French), lightweight, very fast, and excellent quality for its size. Licensed under Apache 2.0.
The first time, Ollama downloads the model (about 3.4 GB in Q4). Subsequent launches are instantaneous. When the >>> prompt appears, you can start chatting.
#4. Verify that the GPU is being used
By default, Ollama detects your NVIDIA GPU and loads the model onto it. To confirm, run this command while a conversation is in progress:
In the PROCESSOR column, you should see 100% GPU. If you see CPU or a partial percentage, the model is spilling over into RAM.
To monitor GPU usage live, open another PowerShell window:
#5. A clean interface with Open WebUI
The terminal is fine for testing. For daily use, Open WebUI looks like ChatGPT—history, Markdown, attachments, everything is there. The simplest approach uses Docker Desktop.
Then open http://localhost:3000 in your browser. Create an account (it is local; nothing leaves your machine), and your Ollama models appear in the list.
#Troubleshooting
- “ollama: command not found”
- The PATH was not reloaded. Close and reopen all your terminals. As a last resort, restart the Windows session.
- Download stuck at 0%
- An enterprise firewall sometimes blocks Ollama. Try from another network or configure a proxy through the HTTPS_PROXY variable.
- The model runs, but it is slow
- Check ollama ps. If it's 100% CPU, your GPU is either too small or not detected correctly. Update your NVIDIA drivers.
- Responses cut off after a few lines
- The default context window is 2048 tokens. Increase it with /set parameter num_ctx 8192 in the conversation.
#Go further
You have Ollama running. The next logical steps are:
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.