what is Ollama and how does it work (guide beginner)
What is Ollama? It is the simplest tool for running an LLM locally on your own machine—without the cloud, without a subscription, and without sending a line to OpenAI. This guide explains from scratch what Ollama is, how it downloads and runs models, the three or four commands you will use 90% of the time, and the bare minimum you need to get started.
#What is Ollama?
Ollama is an open-source tool that downloads, stores, and runs language models (LLMs) directly on your computer. You type a command, the model loads into memory, and you chat with it—either from the command line or through a graphical interface you connect on top. Everything happens locally: no data is sent over the internet once the model has been downloaded.
In practice, Ollama plays two roles at once: it is a daemon (a program that runs in the background) and a CLI (a command-line interface) that communicates with that daemon. The daemon listens on port 11434 on your machine and exposes a small HTTP API—almost the same as OpenAI’s. Any compatible interface (Open WebUI, LM Studio in client mode, Continue.dev, and so on) can therefore connect to it without modification.
Why are so many people switching to it? Because it eliminates all the technical friction of local LLMs: no compilation, no manual GPU configuration, no Python version management. You install it, type ollama run qwen3.5:9b, and it works. It's the equivalent of Docker for LLMs: the model and its runtime environment are packaged in a single format.
#How it works internally
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
When you launch Ollama, three layers work together. Understanding these layers radically changes your ability to diagnose a problem.
- The daemon (ollama serve)
- A service that runs in the background, listens on http://localhost:11434, and manages loading and unloading models in memory. On Windows and macOS, it starts automatically after installation. On Linux, it is typically a systemd service.
- The inference engine (llama.cpp)
- This is what runs the models. Ollama includes a precompiled version with support for CUDA (NVIDIA), Metal (Apple Silicon), and CPU. You do not have to compile anything.
- The CLI (ollama)
- The command-line program you call. It doesn't do much itself: it sends HTTP requests to the daemon and displays the responses.
When you type ollama run qwen3.5:9b, here's what actually happens: the CLI asks the daemon to load this model; if the model isn't on disk, the daemon downloads it from the Ollama registry (a bit like Docker Hub); it loads it into VRAM if you have a compatible GPU, otherwise into RAM; and it opens an interactive chat session. All the heavy lifting happens on the daemon side—the CLI is just a client.
#The commands: run, pull, list
Three commands cover 90% of usage. Learn them, and the rest will seem obvious.
#ollama run
The command to know first. It downloads the model if it is not already present, loads it into memory, and then opens an interactive conversation in the terminal.
Once you’re in the session, type your question and press Enter. To quit, /bye or Ctrl+D. To switch to multiline mode, start with three quotation marks ("""). Anything beginning with / is an internal command (/show, /set, /clear, /save, /load).
You can also pass a prompt directly as an argument—useful for scripting:
#ollama pull
Downloads a model without running it. Useful when you want to prepare several models in advance or pull a specific variant (size, quantization).
The name follows the modele:tag format. Without a tag, you get the default version (usually an 8B/9B with Q4_K_M quantization). With an explicit tag, you choose the size (2b, 4b, 9b, 27b, 30b) and quantization (q4_K_M, q5_K_M, q8_0, fp16).
#ollama list
Displays the models installed on your machine, their disk size, and their installation date.
#Other useful commands
- ollama ps
- Displays the models currently loaded in memory (and how much VRAM/RAM they consume). Essential for diagnosing a "it's sluggish" problem.
- ollama rm <model>
- Delete a model from disk. Useful when the models folder starts weighing 50 GB.
- ollama show <model>
- Displays a model's metadata: context size, quantization, prompt template, capabilities (tools, vision, etc.).
- ollama serve
- Manually start the daemon in the foreground. Useful on Linux or for troubleshooting (you can see the logs live).
- ollama create
- Creates a custom model from a Modelfile—the equivalent of a Dockerfile for LLMs. More advanced topic.
#Where models are stored
Models can weigh several gigabytes. Knowing where they live is essential for cleaning up, moving the library to another drive, or backing it up.
- Windows
- %USERPROFILE%\.ollama\models — typically C:\Users\yourname\.ollama\models
- macOS
- ~/.ollama/models
- Linux (official script installation)
- /usr/share/ollama/.ollama/models — le démon tourne sous l'utilisateur ollama
- Linux (manual user installation)
- ~/.ollama/models
Inside this folder are two subfolders: blobs (the binary model files, identified by a hash) and manifests (the metadata describing which blob corresponds to which tag). Do not edit these files manually—use ollama rm et ollama pull to manage them.
#Minimal configuration to get started
The question "what is Ollama and can my machine run it?" comes up constantly. Here are the real thresholds, not the marketing minimums.
- 8 GB RAM, no GPU
- Limited to 2–3B models in Q4 (Qwen 3.5 2B, Granite 4.2 3B). 5–10 tokens/second on a modern CPU. Usable for testing and learning.
- 16 GB of RAM, no GPU
- 8B models fit in Q4 (Granite 4.2 8B, Qwen 3.5 9B). Expect 4–8 tokens/second. Too slow for sustained interactive use, but it works.
- GPU with 8 GB of VRAM
- Comfortable with 8B models in Q4 (Qwen 3.5 9B, Granite 4.2 8B — 30–50 tokens/s). The limit is reached with 12–14B models.
- 12 GB VRAM GPU (RTX 3060, 4070)
- The 2026 entry-level sweet spot. All the smooth 8B models, Gemma 4 12B in comfortable Q4, or Qwen 3.5 9B in Q8 for the best quality in this tier.
- 16 GB GPU (RTX 4080, 5080) or 32 GB unified-memory Mac M-Pro
- We move up to 24B models in Q4 (Mistral Small 24B, Devstral 24B, gpt-oss 20B) or to Qwen 3.5 9B in Q8 for quality.
- 24 GB GPU (RTX 3090, 4090) or 48 GB+ Mac M-Max
- You reach the comfortable 27–35B models (Qwen 3.8 27B, Qwen 3.6 35B-A3B) and large Q4 MoE models with a little patience.
For storage, plan on 10 GB per 7B model in Q4 and 30 GB per 32B model. Having 50 to 100 GB free from the start prevents you from having to clean up too soon.
#Recommended first model
When you discover Ollama, the instinct is to test the largest model possible “just to see.” That is a mistake—you saturate the VRAM, it swaps to RAM, it is slow, and you give up. Start small.
- 01Qwen 3.5 9B — the reasonable defaultollama run qwen3.5:9b downloads the Q4_K_M version of Qwen 3.5 9B (6.6 GB). Good overall quality, strong French, vision input, 256k context window. It’s the 8 GB choice for 2026 and an excellent benchmark for evaluating what your machine can do.
- 02Granite 4.2 8B — the no-frills alternativeollama run granite4.2:8b for IBM’s model (5.3 GB, Apache 2.0). Slightly lighter, highly token-efficient, with a 128k context. A good compromise if your card is tight on VRAM. For very natural French with more VRAM, Mistral Small 24B (ollama run mistral-small) remains a safe bet.
- 03Qwen 3.5 4B — the rising all-rounderollama run qwen3.5:4b for the new default small model (3.4 GB, Apache 2.0). Surprisingly capable for its size, adequate for light coding and extraction. Many adopt it as their default when VRAM is limited.
- 04Granite 4.2 3B — if your machine is modestollama run granite4.2:3b for a model that fits in 4 GB of RAM (2.2 GB, Apache 2.0). Less capable but usable for simple tasks and for learning how Ollama works without suffering through long loading times.
#Tips and pitfalls
- The model unloads after 5 minutes of inactivity
- This is OLLAMA_KEEP_ALIVE's default value. If you want to keep it in VRAM longer, set OLLAMA_KEEP_ALIVE=2h (or -1 to never unload it). Conversely, set it to 0 if you want to free the VRAM as soon as a request finishes.
- Responses truncated after a few sentences
- The default num_ctx parameter (context window) is 4096 tokens, regardless of what the model supports. For RAG use or long documents, create a Modelfile with PARAMETER num_ctx 32768 and run ollama create. Be aware that VRAM usage increases quickly.
- No shell autocompletion
- Type ollama help to see all commands, and ollama help <command> for details about each one. The docs are in the CLI.
- Connection refused from another machine on the network
- By default, Ollama listens only on 127.0.0.1. To expose it on the LAN, set OLLAMA_HOST=0.0.0.0:11434 before starting the daemon. Remember the firewall.
- One model name, multiple versions
- ollama pull qwen3.5:9b and ollama pull qwen3.5:9b-q8_0 download two separate variants. Monitor your disk with ollama list.
#Go further
You understand what Ollama is and how it basically works. The natural next directions are:
- Actually install it on your OS
- The guide Install Ollama on Windows (or macOS or Linux) walks you through the installation, GPU verification, and first model in a clean environment.
- Choosing the right quantization
- The Choosing Your Quantization guide (Q4, Q5, Q8, FP16) visually compares the real quality losses and helps you weigh quality against VRAM.
- Connect a real graphical interface
- The Open WebUI guide with Ollama installs a multi-user ChatGPT-like interface with built-in RAG in 10 minutes, on top of the daemon you now have.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.