Beginner 7 minOllama

what is Ollama and how does it work (guide beginner)

What is Ollama? It is the simplest tool for running an LLM locally on your own machine—without the cloud, without a subscription, and without sending a line to OpenAI. This guide explains from scratch what Ollama is, how it downloads and runs models, the three or four commands you will use 90% of the time, and the bare minimum you need to get started.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#What is Ollama?

Ollama is an open-source tool that downloads, stores, and runs language models (LLMs) directly on your computer. You type a command, the model loads into memory, and you chat with it—either from the command line or through a graphical interface you connect on top. Everything happens locally: no data is sent over the internet once the model has been downloaded.

In practice, Ollama plays two roles at once: it is a daemon (a program that runs in the background) and a CLI (a command-line interface) that communicates with that daemon. The daemon listens on port 11434 on your machine and exposes a small HTTP API—almost the same as OpenAI’s. Any compatible interface (Open WebUI, LM Studio in client mode, Continue.dev, and so on) can therefore connect to it without modification.

i
In one sentence
Ollama = a local daemon that can download and run open-weight LLMs, plus a CLI for communicating with it. Under the hood, it is llama.cpp packaged cleanly. You will not need to compile anything or handle GGUF files manually.

Why are so many people switching to it? Because it eliminates all the technical friction of local LLMs: no compilation, no manual GPU configuration, no Python version management. You install it, type ollama run qwen3.5:9b, and it works. It's the equivalent of Docker for LLMs: the model and its runtime environment are packaged in a single format.

#How it works internally

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

When you launch Ollama, three layers work together. Understanding these layers radically changes your ability to diagnose a problem.

The daemon (ollama serve)
A service that runs in the background, listens on http://localhost:11434, and manages loading and unloading models in memory. On Windows and macOS, it starts automatically after installation. On Linux, it is typically a systemd service.
The inference engine (llama.cpp)
This is what runs the models. Ollama includes a precompiled version with support for CUDA (NVIDIA), Metal (Apple Silicon), and CPU. You do not have to compile anything.
The CLI (ollama)
The command-line program you call. It doesn't do much itself: it sends HTTP requests to the daemon and displays the responses.

When you type ollama run qwen3.5:9b, here's what actually happens: the CLI asks the daemon to load this model; if the model isn't on disk, the daemon downloads it from the Ollama registry (a bit like Docker Hub); it loads it into VRAM if you have a compatible GPU, otherwise into RAM; and it opens an interactive chat session. All the heavy lifting happens on the daemon side—the CLI is just a client.

→
Port 11434
This is the HTTP port everyone uses. If an app says it is “connecting to your Ollama,” it is talking to http://localhost:11434. If Ollama is running and you open this URL in a browser, you’ll see the message Ollama is running. This is the simplest health check.

#The commands: run, pull, list

Three commands cover 90% of usage. Learn them, and the rest will seem obvious.

#ollama run

The command to know first. It downloads the model if it is not already present, loads it into memory, and then opens an interactive conversation in the terminal.

Terminal
ollama run qwen3.5:9b
# Première fois : télécharge ~6.6 Go puis lance le chat
# Fois suivantes : charge le modèle déjà présent et démarre

Once you’re in the session, type your question and press Enter. To quit, /bye or Ctrl+D. To switch to multiline mode, start with three quotation marks ("""). Anything beginning with / is an internal command (/show, /set, /clear, /save, /load).

You can also pass a prompt directly as an argument—useful for scripting:

One-shot prompt
ollama run qwen3.5:9b "Résume ce texte en 3 points : [...]"

#ollama pull

Downloads a model without running it. Useful when you want to prepare several models in advance or pull a specific variant (size, quantization).

Terminal
ollama pull granite4.2:3b
ollama pull qwen3.5:9b
ollama pull qwen3.5:9b-q8_0

The name follows the modele:tag format. Without a tag, you get the default version (usually an 8B/9B with Q4_K_M quantization). With an explicit tag, you choose the size (2b, 4b, 9b, 27b, 30b) and quantization (q4_K_M, q5_K_M, q8_0, fp16).

i
Quantization, in a nutshell
A "raw" model weighs 14 GB for 7 billion parameters in FP16. Quantization compresses these numbers into fewer bits: Q4_K_M (4 bits) divides the size by ~4 with minimal quality loss. This is the recommended sweet spot for getting started. Q5_K_M offers slightly better quality, Q8_0 is nearly perfect but 2x heavier, and FP16 is reserved for powerful GPUs.

#ollama list

Displays the models installed on your machine, their disk size, and their installation date.

Terminal
ollama list
# NAME               ID            SIZE    MODIFIED
# qwen3.5:9b         42182419e950  6.6 GB  2 days ago
# granite4.2:8b      f974a74358d6  5.3 GB  1 week ago

#Other useful commands

ollama ps
Displays the models currently loaded in memory (and how much VRAM/RAM they consume). Essential for diagnosing a "it's sluggish" problem.
ollama rm <model>
Delete a model from disk. Useful when the models folder starts weighing 50 GB.
ollama show <model>
Displays a model's metadata: context size, quantization, prompt template, capabilities (tools, vision, etc.).
ollama serve
Manually start the daemon in the foreground. Useful on Linux or for troubleshooting (you can see the logs live).
ollama create
Creates a custom model from a Modelfile—the equivalent of a Dockerfile for LLMs. More advanced topic.

#Where models are stored

Models can weigh several gigabytes. Knowing where they live is essential for cleaning up, moving the library to another drive, or backing it up.

Windows
%USERPROFILE%\.ollama\models — typically C:\Users\yourname\.ollama\models
macOS
~/.ollama/models
Linux (official script installation)
/usr/share/ollama/.ollama/models — le démon tourne sous l'utilisateur ollama
Linux (manual user installation)
~/.ollama/models

Inside this folder are two subfolders: blobs (the binary model files, identified by a hash) and manifests (the metadata describing which blob corresponds to which tag). Do not edit these files manually—use ollama rm et ollama pull to manage them.

→
Move models to another drive
If your system SSD is small and you have a larger secondary drive, set the OLLAMA_MODELS environment variable to point elsewhere. Example: OLLAMA_MODELS=D:\ollama-models on Windows, or export OLLAMA_MODELS=/mnt/data/ollama-models on Linux. Restart the daemon for the variable to take effect.

#Minimal configuration to get started

The question "what is Ollama and can my machine run it?" comes up constantly. Here are the real thresholds, not the marketing minimums.

8 GB RAM, no GPU
Limited to 2–3B models in Q4 (Qwen 3.5 2B, Granite 4.2 3B). 5–10 tokens/second on a modern CPU. Usable for testing and learning.
16 GB of RAM, no GPU
8B models fit in Q4 (Granite 4.2 8B, Qwen 3.5 9B). Expect 4–8 tokens/second. Too slow for sustained interactive use, but it works.
GPU with 8 GB of VRAM
Comfortable with 8B models in Q4 (Qwen 3.5 9B, Granite 4.2 8B — 30–50 tokens/s). The limit is reached with 12–14B models.
12 GB VRAM GPU (RTX 3060, 4070)
The 2026 entry-level sweet spot. All the smooth 8B models, Gemma 4 12B in comfortable Q4, or Qwen 3.5 9B in Q8 for the best quality in this tier.
16 GB GPU (RTX 4080, 5080) or 32 GB unified-memory Mac M-Pro
We move up to 24B models in Q4 (Mistral Small 24B, Devstral 24B, gpt-oss 20B) or to Qwen 3.5 9B in Q8 for quality.
24 GB GPU (RTX 3090, 4090) or 48 GB+ Mac M-Max
You reach the comfortable 27–35B models (Qwen 3.8 27B, Qwen 3.6 35B-A3B) and large Q4 MoE models with a little patience.

For storage, plan on 10 GB per 7B model in Q4 and 30 GB per 32B model. Having 50 to 100 GB free from the start prevents you from having to clean up too soon.

!
The “GPU detected” but unused trap
On Windows and Linux, Ollama detects the GPU automatically—unless your NVIDIA drivers are too old or you have an AMD GPU without ROCm properly installed. Check with ollama ps: if the PROCESSOR column shows 100% CPU while you have a GPU, hardware acceleration is not active. On Apple Silicon, Metal acceleration is automatic and always active.

#Recommended first model

When you discover Ollama, the instinct is to test the largest model possible “just to see.” That is a mistake—you saturate the VRAM, it swaps to RAM, it is slow, and you give up. Start small.

  1. 01
    Qwen 3.5 9B — the reasonable default
    ollama run qwen3.5:9b downloads the Q4_K_M version of Qwen 3.5 9B (6.6 GB). Good overall quality, strong French, vision input, 256k context window. It’s the 8 GB choice for 2026 and an excellent benchmark for evaluating what your machine can do.
  2. 02
    Granite 4.2 8B — the no-frills alternative
    ollama run granite4.2:8b for IBM’s model (5.3 GB, Apache 2.0). Slightly lighter, highly token-efficient, with a 128k context. A good compromise if your card is tight on VRAM. For very natural French with more VRAM, Mistral Small 24B (ollama run mistral-small) remains a safe bet.
  3. 03
    Qwen 3.5 4B — the rising all-rounder
    ollama run qwen3.5:4b for the new default small model (3.4 GB, Apache 2.0). Surprisingly capable for its size, adequate for light coding and extraction. Many adopt it as their default when VRAM is limited.
  4. 04
    Granite 4.2 3B — if your machine is modest
    ollama run granite4.2:3b for a model that fits in 4 GB of RAM (2.2 GB, Apache 2.0). Less capable but usable for simple tasks and for learning how Ollama works without suffering through long loading times.
→
Run ollama ps during a chat
In another terminal while you chat, run ollama ps. You’ll see the loaded model, memory usage, and GPU/CPU split. It’s the best way to understand what’s actually happening.

#Tips and pitfalls

The model unloads after 5 minutes of inactivity
This is OLLAMA_KEEP_ALIVE's default value. If you want to keep it in VRAM longer, set OLLAMA_KEEP_ALIVE=2h (or -1 to never unload it). Conversely, set it to 0 if you want to free the VRAM as soon as a request finishes.
Responses truncated after a few sentences
The default num_ctx parameter (context window) is 4096 tokens, regardless of what the model supports. For RAG use or long documents, create a Modelfile with PARAMETER num_ctx 32768 and run ollama create. Be aware that VRAM usage increases quickly.
No shell autocompletion
Type ollama help to see all commands, and ollama help <command> for details about each one. The docs are in the CLI.
Connection refused from another machine on the network
By default, Ollama listens only on 127.0.0.1. To expose it on the LAN, set OLLAMA_HOST=0.0.0.0:11434 before starting the daemon. Remember the firewall.
One model name, multiple versions
ollama pull qwen3.5:9b and ollama pull qwen3.5:9b-q8_0 download two separate variants. Monitor your disk with ollama list.
i
A graphical interface in 5 minutes
The CLI is enough to understand the basics, but for daily use, connect Open WebUI (a ChatGPT-style web interface) or LM Studio in client mode on top of the Ollama daemon. These tools automatically detect the server at localhost:11434 and expose all your models in a clean interface.

#Go further

You understand what Ollama is and how it basically works. The natural next directions are:

Actually install it on your OS
The guide Install Ollama on Windows (or macOS or Linux) walks you through the installation, GPU verification, and first model in a clean environment.
Choosing the right quantization
The Choosing Your Quantization guide (Q4, Q5, Q8, FP16) visually compares the real quality losses and helps you weigh quality against VRAM.
Connect a real graphical interface
The Open WebUI guide with Ollama installs a multi-user ChatGPT-like interface with built-in RAG in 10 minutes, on top of the daemon you now have.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.