Beginner 3 minOllama

Install Ollama on Windows 11: complete guide (2026)

Direct response

On Windows 11, Ollama installs through a standard installer downloaded from the official website, with automatic GPU detection (native CUDA NVIDIA support). A single terminal command (ollama run qwen3.5:4b) downloads and launches a first model in about 5 minutes, with no manual configuration or complex command line.

You have a Windows PC and want to run an LLM locally without using ChatGPT. Ollama is probably the simplest tool for the job: a standard installer, one command, one model. In exactly 3 minutes, you’ll have Qwen 3.5 running on your machine.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows 11

#Why Ollama?

Ollama handles all the tedious work for you: downloading models, quantization, loading into VRAM, GPU detection, and the local HTTP API. Under the hood, it’s llama.cpp. But you won’t have to compile it.

The tool is free, open source, and runs entirely offline once the model has been downloaded. No data is sent to a remote server.

i
In two words
Ollama = a local daemon running in the background + a CLI for communicating with it. The models live in a folder on your disk. That's it.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Windows 10 or 11
64-bit. Ollama has no 32-bit version.
8 GB of RAM minimum
16 GB is comfortable for 8B to 9B models in Q4 quantization.
10 GB of disk space
A 4B to 9B model weighs about 3 to 7 GB. Leave plenty of room if you plan to test several.
GPU NVIDIA (optional)
Up-to-date drivers, RTX 20/30/40/50 or GTX 16. Ollama detects the GPU automatically.
→
It also works without a GPU
A modern CPU can run a small model (2 to 4B, such as Qwen 3.5 2B or Granite 4.2 3B) at 5–10 tokens/second. It’s slow but usable for testing. For daily use, target a GPU with at least 6 GB of VRAM.

#1. Download

Go to the official website and download the Windows installer.

Official link
https://ollama.com/download/windows

The installer is about 700 MB. That's normal: it includes precompiled llama.cpp binaries with CUDA support.

!
Check the domain
Download Ollama only from ollama.com. Versions redistributed on GitHub by third parties may be compromised. The tool is open source—if you are paranoid, compile it from the official repo.

#2. Installation

  1. 01
    Run OllamaSetup.exe
    Double-click the downloaded installer. Windows Defender may display a SmartScreen warning — click "More info," then "Run anyway."
  2. 02
    Silent installation
    No configuration window: the installer installs to %LOCALAPPDATA%\Programs\Ollama and automatically starts the service in the background.
  3. 03
    Check the system tray icon
    A small Ollama icon appears near the clock in the bottom-right corner. It indicates that the daemon is running.
  4. 04
    Open a terminal
    PowerShell, Windows Terminal, or cmd.exe—it doesn't matter. Type the verification command below.
PowerShell
ollama --version

You should see something like ollama version is 0.5.x. If the command isn't recognized, close and reopen the terminal—the PATH was just modified.

#3. Your first model

Qwen 3.5 4B is a good first choice: multilingual (including very good French), lightweight, very fast, and excellent quality for its size. Licensed under Apache 2.0.

Download and run
ollama run qwen3.5:4b

The first time, Ollama downloads the model (about 3.4 GB in Q4). Subsequent launches are instantaneous. When the >>> prompt appears, you can start chatting.

Conversation trial
>>> Bonjour, présente-toi en 2 phrases.

Je suis Qwen 3.5, un modèle de langage open-source développé par l'équipe Qwen (Alibaba).
Je fonctionne entièrement en local sur votre machine, sans connexion internet.
→
Leave the conversation
Type /bye or Ctrl+D to exit. The model remains loaded in memory for a few minutes, making restarts instantaneous.

#4. Verify that the GPU is being used

By default, Ollama detects your NVIDIA GPU and loads the model onto it. To confirm, run this command while a conversation is in progress:

Loaded model status
ollama ps

In the PROCESSOR column, you should see 100% GPU. If you see CPU or a partial percentage, the model is spilling over into RAM.

i
Not enough VRAM?
Qwen 3.5 4B Q4 needs about 3.4 GB of VRAM. If your GPU is really tight on memory, switch to an even smaller model: Granite 4.2 3B (ollama run granite4.2:3b, 2.2 GB) or Qwen 3.5 2B (ollama run qwen3.5:2b, 1.9 GB).

To monitor GPU usage live, open another PowerShell window:

Monitoring
nvidia-smi -l 1

#5. A clean interface with Open WebUI

The terminal is fine for testing. For daily use, Open WebUI looks like ChatGPT—history, Markdown, attachments, everything is there. The simplest approach uses Docker Desktop.

Run Open WebUI
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main

Then open http://localhost:3000 in your browser. Create an account (it is local; nothing leaves your machine), and your Ollama models appear in the list.

#Troubleshooting

“ollama: command not found”
The PATH was not reloaded. Close and reopen all your terminals. As a last resort, restart the Windows session.
Download stuck at 0%
An enterprise firewall sometimes blocks Ollama. Try from another network or configure a proxy through the HTTPS_PROXY variable.
The model runs, but it is slow
Check ollama ps. If it's 100% CPU, your GPU is either too small or not detected correctly. Update your NVIDIA drivers.
Responses cut off after a few lines
The default context window is 2048 tokens. Increase it with /set parameter num_ctx 8192 in the conversation.

#Go further

You have Ollama running. The next logical steps are:

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.