Beginner 11 minConfiguration

Understanding the window of context

Direct response

The context window is the maximum number of tokens the model processes at once: the system prompt, history, attached documents, and response currently being generated, all added together. It consumes memory because each token keeps an entry in the KV cache. Under Ollama, the default is 4,096 tokens on a card with less than 24 GiB of VRAM: often, it—not the model—is what limits what you can have it read.

If the window is too short, it forgets the beginning of a conversation or truncates a document. If it is too large, it saturates the VRAM and slows everything down. This guide explains what it contains, calculates its memory usage from the architecture of a real model, and shows how to tune it without unpleasant surprises.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What the context window contains

According to Ollama's documentation, context length is the maximum number of tokens the model can access in memory. Everything counts toward this single budget: the system message, conversation history, files or document passages you paste, your latest question, and the response the model is currently writing. For a reasoning model, thinking tokens count too. When the total exceeds the window, something has to go: tools generally truncate the beginning or reject the request. The model has no memory outside this window unless an external system (summary, RAG) feeds information back into it.

i
Window and memory are not the same thing
An assistant that “remembers” yesterday’s conversation does not do so because of the context window: the application rereads a history or database and copies it into the prompt. The window is the limit of what can fit in it at any given moment.

#The token: the unit that fills the window

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A token is not a word: it is a fragment of text defined by the model’s tokenizer, often a syllable or a common word. A rare or long word may use several. French generally consumes more tokens than English for equivalent content because most tokenizers are trained primarily on English. The exact ratio varies from model to model: rather than relying on a rule of thumb, measure it. The Ollama API returns prompt_eval_count, the number of prompt tokens, with every request.

Counting a prompt's actual tokens
curl http://localhost:11434/api/generate -d '{"model":"qwen3:8b","prompt":"Bonjour le monde","stream":false}'
# lisez prompt_eval_count et eval_count dans la réponse JSON
Rule of thumb
A page of dense French text in A4 format contains roughly one thousand tokens, depending on the tokenizer: validate this with prompt_eval_count on your own document.
A book
Several hundred thousand tokens: beyond the reach of a 32,000- or 64,000-token window without chunking.
One response
A reasoning model can produce thousands of reasoning tokens before the visible response, consuming the context window.

#What context window size should you choose?

Ollama sets its default based on video memory: approximately 4 000 tokens (4k) below 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k starting at 48 GiB. The same page recommends at least 64 000 tokens for tasks that require a large context, such as web research, agents, and coding tools. Recent models advertise much higher maximums: the QuelLLM catalog lists, for example, approximately 256 000 tokens for Kimi K2.5 and approximately 1 million for Kimi K3, DeepSeek V4 Flash, and GLM 5.2. But an advertised maximum is not a usable context on your machine: memory, and sometimes quality, stand in the way.

Which context window for which use case
UsageIndicative windowNote
Short conversation, question-and-answer4,096 to 8,192 tokensEnough if you don't paste documents
Summary or analysis of an article16,000 to 32,000 tokensCheck the text's token count
Code assistant for a repository64,000 tokens or moreRecommended by Ollama for coding tools
Agent with tools and web search64,000 tokens or moreEach tool call reinjects text
Very large corpusDon't target the windowUse RAG instead of sending everything

Two sizing pitfalls come up often. First, the window must contain the response: if you fill 31,000 of 32,000 tokens with a document, there is almost nothing left to answer with, and a reasoning model will stop mid-thought. Second, in a conversation, the history grows with every turn: a window sufficient for the first message may be full by the twentieth. So plan for the worst case of your usage, not the average case, and keep a margin of about one-fifth of the window.

#How much memory does context consume?

Each token in the context window leaves a key and a value in every model layer, stored in the KV cache. The per-token cost is calculated from four architectural values: the number of layers, the number of key-value heads, the head dimension, and the size of a number (2 bytes in FP16). The formula is: 2 (key and value) × layers × KV heads × head dimension × 2 bytes. Note: the key-value heads matter, not the attention heads, because recent models share several of them (grouped-query attention). Using the attention heads overestimates the cache by a factor of four for the model below.

Take Qwen3-8B, whose official specifications list 36 layers and 8 key-value heads (versus 32 query heads), with the public configuration setting the head dimension to 128. The cost is 2 × 36 × 8 × 128 × 2 = 147,456 bytes per token, or 144 KiB.

Qwen3-8B KV cache in FP16 (calculated from its configuration)
ContextKV cache (FP16)KV cache (q8_0, approximately half)
4,096 tokens0.56 GiB0.28 GiB
8,192 tokens1.13 GiB0.56 GiB
16,384 tokens2.25 GiB1.13 GiB
32 768 tokens4.50 GiB2.25 GiB
131,072 tokens (with YaRN, according to the specifications)18.00 GiB9.00 GiB
!
The 8 GB card trap
Qwen3-8B in Q4 uses about 5 GB for its weights. With a 32,768-token context, the KV cache adds 4.5 GiB, bringing the total to about 9.5 GB even before compute buffers. On an 8 GB card, the model is then partially offloaded to the CPU and generation slows dramatically, without a clear error message. Check with ollama ps.

Two levers reduce cache usage. The first is cache quantization: the Ollama FAQ states that q8_0 uses about half the memory of FP16 with very little loss, while q4_0 uses about a quarter with more noticeable loss at larger contexts; they require Flash Attention to be enabled. The second is reducing the window to what you need. Our KV cache guide details the settings.

#Adjust the context length

#Under Ollama

Ollama lets you set the default length when the server starts, change it for a session, or specify it per request through the API.

Three ways to set the context in Ollama
# Valeur par défaut pour tout le serveur
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

# Pour une session interactive
>>> /set parameter num_ctx 16384

# Pour un modèle personnalisé (Modelfile)
FROM qwen3:8b
PARAMETER num_ctx 16384

After loading, run ollama ps: the CONTEXT column displays the allocated length, and the PROCESSOR column shows the split between GPU and CPU. If part of the workload runs on the CPU, reduce the context or choose a smaller model.

#Under LM Studio

In LM Studio, the context length is set when loading the model, in the load settings. Changing the value requires reloading the model. Check the memory estimate shown before confirming.

#A four-step tuning procedure

  1. 01
    Measure the actual need
    Send your typical document or history and read prompt_eval_count. Add the expected response length, along with reasoning if the model produces it.
  2. 02
    Choose the window
    Choose the smallest value that contains this total with a 20% margin, at least 64,000 tokens if you use an agent or coding tool, as Ollama recommends.
  3. 03
    Check the memory
    Load the model with this context window and run ollama ps. The processor must show 100% GPU; otherwise, reduce the window, quantize the cache, or change models.
  4. 04
    Test recall
    Place a specific fact in the middle of a long text and ask for it. If the model misses it, the advertised context window exceeds what it can use, and chunking or RAG is required.

#A large context window doesn't mean faithful reading

A 2023 study, “Lost in the Middle,” shows that model performance can deteriorate significantly depending on where information appears in the context: it is often better when the information is at the beginning or end, and deteriorates when it is in the middle, even for models advertised as supporting long contexts. The RULER benchmark, published in 2024, goes further: almost all tested models lose substantial accuracy as length increases, and only half of them maintain a satisfactory level at 32,000 tokens, even though all advertised 32,000 or more.

These studies concern models from their respective eras and say nothing about 2026 models, several of which are trained specifically for long context. But the guidance still holds: put the prompt and critical facts at the beginning, repeat the question at the end, and measure performance on your documents before trusting a claimed window of several hundred thousand tokens. The simplest test is to place a precise piece of information in the middle of a long text from your domain, then ask for it again: if the model retrieves it every time, the window is usable for your purposes; if it misses it, shorten the submitted text or use chunking.

#When the content overflows

Summarize as you go
Have it produce a summary of the exchange every few turns, then continue with that summary in the context: you lose detail, but keep the thread.
Split the document
A long PDF is processed section by section, then the partial answers are aggregated. The guide to chunking details useful sizes.
Switch to RAG
When the corpus far exceeds the context window, retrieve the few relevant passages instead of sending everything.
Choose a model with a larger context window
Only if your memory can handle it: refer to the KV cache calculation above before doubling the window.
FAQ
What is an LLM's context window?+
This is the maximum number of tokens the model processes at once: system message, history, attached documents, current question, and response. Beyond that, the beginning is truncated or the request is rejected. The window limits what the model can read, not what it learned during training.
What is the default context window of Ollama?+
It depends on video memory: 4k tokens under 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above 48 GiB. Ollama recommends at least 64,000 tokens for agents, web research, and coding tools. You can change it with OLLAMA_CONTEXT_LENGTH or num_ctx.
How can you increase a local model's context?+
Set OLLAMA_CONTEXT_LENGTH when starting the server, use /set parameter num_ctx in an interactive session, or declare PARAMETER num_ctx in a Modelfile. Then check with ollama ps that the model remains entirely on the GPU: a longer context uses more memory and may spill over to the CPU.
How much VRAM does a 32,000-token context consume?+
It depends on the architecture. For Qwen3-8B, the KV cache in FP16 reaches about 4.5 GiB at 32,768 tokens, in addition to the weights. The formula is: 2 × layers × KV heads × head dimension × 2 bytes per token. q8_0 cache quantization cuts this cost roughly in half.
Does a larger context make the model smarter?+
No. It lets it read more text, not reason better. On the contrary, the “Lost in the Middle” and RULER studies show that accuracy can drop as the context grows. Use the window your task requires, not the largest one possible.
Do you need RAG or a large context window to analyze documents?+
If the document fits within the window with room to spare and you have the memory, sending it in full is simpler. Beyond that, or when querying a database containing many files, RAG is more economical and often more reliable because it sends only the relevant passages.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.