Understanding the window of context
The context window is the maximum number of tokens the model processes at once: the system prompt, history, attached documents, and response currently being generated, all added together. It consumes memory because each token keeps an entry in the KV cache. Under Ollama, the default is 4,096 tokens on a card with less than 24 GiB of VRAM: often, it—not the model—is what limits what you can have it read.
If the window is too short, it forgets the beginning of a conversation or truncates a document. If it is too large, it saturates the VRAM and slows everything down. This guide explains what it contains, calculates its memory usage from the architecture of a real model, and shows how to tune it without unpleasant surprises.
#What the context window contains
According to Ollama's documentation, context length is the maximum number of tokens the model can access in memory. Everything counts toward this single budget: the system message, conversation history, files or document passages you paste, your latest question, and the response the model is currently writing. For a reasoning model, thinking tokens count too. When the total exceeds the window, something has to go: tools generally truncate the beginning or reject the request. The model has no memory outside this window unless an external system (summary, RAG) feeds information back into it.
#The token: the unit that fills the window
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A token is not a word: it is a fragment of text defined by the model’s tokenizer, often a syllable or a common word. A rare or long word may use several. French generally consumes more tokens than English for equivalent content because most tokenizers are trained primarily on English. The exact ratio varies from model to model: rather than relying on a rule of thumb, measure it. The Ollama API returns prompt_eval_count, the number of prompt tokens, with every request.
- Rule of thumb
- A page of dense French text in A4 format contains roughly one thousand tokens, depending on the tokenizer: validate this with prompt_eval_count on your own document.
- A book
- Several hundred thousand tokens: beyond the reach of a 32,000- or 64,000-token window without chunking.
- One response
- A reasoning model can produce thousands of reasoning tokens before the visible response, consuming the context window.
#What context window size should you choose?
Ollama sets its default based on video memory: approximately 4 000 tokens (4k) below 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k starting at 48 GiB. The same page recommends at least 64 000 tokens for tasks that require a large context, such as web research, agents, and coding tools. Recent models advertise much higher maximums: the QuelLLM catalog lists, for example, approximately 256 000 tokens for Kimi K2.5 and approximately 1 million for Kimi K3, DeepSeek V4 Flash, and GLM 5.2. But an advertised maximum is not a usable context on your machine: memory, and sometimes quality, stand in the way.
| Usage | Indicative window | Note |
|---|---|---|
| Short conversation, question-and-answer | 4,096 to 8,192 tokens | Enough if you don't paste documents |
| Summary or analysis of an article | 16,000 to 32,000 tokens | Check the text's token count |
| Code assistant for a repository | 64,000 tokens or more | Recommended by Ollama for coding tools |
| Agent with tools and web search | 64,000 tokens or more | Each tool call reinjects text |
| Very large corpus | Don't target the window | Use RAG instead of sending everything |
Two sizing pitfalls come up often. First, the window must contain the response: if you fill 31,000 of 32,000 tokens with a document, there is almost nothing left to answer with, and a reasoning model will stop mid-thought. Second, in a conversation, the history grows with every turn: a window sufficient for the first message may be full by the twentieth. So plan for the worst case of your usage, not the average case, and keep a margin of about one-fifth of the window.
#How much memory does context consume?
Each token in the context window leaves a key and a value in every model layer, stored in the KV cache. The per-token cost is calculated from four architectural values: the number of layers, the number of key-value heads, the head dimension, and the size of a number (2 bytes in FP16). The formula is: 2 (key and value) × layers × KV heads × head dimension × 2 bytes. Note: the key-value heads matter, not the attention heads, because recent models share several of them (grouped-query attention). Using the attention heads overestimates the cache by a factor of four for the model below.
Take Qwen3-8B, whose official specifications list 36 layers and 8 key-value heads (versus 32 query heads), with the public configuration setting the head dimension to 128. The cost is 2 × 36 × 8 × 128 × 2 = 147,456 bytes per token, or 144 KiB.
| Context | KV cache (FP16) | KV cache (q8_0, approximately half) |
|---|---|---|
| 4,096 tokens | 0.56 GiB | 0.28 GiB |
| 8,192 tokens | 1.13 GiB | 0.56 GiB |
| 16,384 tokens | 2.25 GiB | 1.13 GiB |
| 32 768 tokens | 4.50 GiB | 2.25 GiB |
| 131,072 tokens (with YaRN, according to the specifications) | 18.00 GiB | 9.00 GiB |
Two levers reduce cache usage. The first is cache quantization: the Ollama FAQ states that q8_0 uses about half the memory of FP16 with very little loss, while q4_0 uses about a quarter with more noticeable loss at larger contexts; they require Flash Attention to be enabled. The second is reducing the window to what you need. Our KV cache guide details the settings.
#Adjust the context length
#Under Ollama
Ollama lets you set the default length when the server starts, change it for a session, or specify it per request through the API.
After loading, run ollama ps: the CONTEXT column displays the allocated length, and the PROCESSOR column shows the split between GPU and CPU. If part of the workload runs on the CPU, reduce the context or choose a smaller model.
#Under LM Studio
In LM Studio, the context length is set when loading the model, in the load settings. Changing the value requires reloading the model. Check the memory estimate shown before confirming.
#A four-step tuning procedure
- 01Measure the actual needSend your typical document or history and read prompt_eval_count. Add the expected response length, along with reasoning if the model produces it.
- 02Choose the windowChoose the smallest value that contains this total with a 20% margin, at least 64,000 tokens if you use an agent or coding tool, as Ollama recommends.
- 03Check the memoryLoad the model with this context window and run ollama ps. The processor must show 100% GPU; otherwise, reduce the window, quantize the cache, or change models.
- 04Test recallPlace a specific fact in the middle of a long text and ask for it. If the model misses it, the advertised context window exceeds what it can use, and chunking or RAG is required.
#A large context window doesn't mean faithful reading
A 2023 study, “Lost in the Middle,” shows that model performance can deteriorate significantly depending on where information appears in the context: it is often better when the information is at the beginning or end, and deteriorates when it is in the middle, even for models advertised as supporting long contexts. The RULER benchmark, published in 2024, goes further: almost all tested models lose substantial accuracy as length increases, and only half of them maintain a satisfactory level at 32,000 tokens, even though all advertised 32,000 or more.
These studies concern models from their respective eras and say nothing about 2026 models, several of which are trained specifically for long context. But the guidance still holds: put the prompt and critical facts at the beginning, repeat the question at the end, and measure performance on your documents before trusting a claimed window of several hundred thousand tokens. The simplest test is to place a precise piece of information in the middle of a long text from your domain, then ask for it again: if the model retrieves it every time, the window is usable for your purposes; if it misses it, shorten the submitted text or use chunking.
#When the content overflows
- Summarize as you go
- Have it produce a summary of the exchange every few turns, then continue with that summary in the context: you lose detail, but keep the thread.
- Split the document
- A long PDF is processed section by section, then the partial answers are aggregated. The guide to chunking details useful sizes.
- Switch to RAG
- When the corpus far exceeds the context window, retrieve the few relevant passages instead of sending everything.
- Choose a model with a larger context window
- Only if your memory can handle it: refer to the KV cache calculation above before doubling the window.
What is an LLM's context window?+
What is the default context window of Ollama?+
How can you increase a local model's context?+
How much VRAM does a 32,000-token context consume?+
Does a larger context make the model smarter?+
Do you need RAG or a large context window to analyze documents?+
#Go further
- Quantize the KV cache: save VRAM
- Tokens and tokenization: understanding what an LLM consumes
- Choose your quantization (Q4, Q5, Q8, FP16)
- What is RAG and how does it work?
- Chunking strategies
- VRAM Calculator
- Source: Ollama documentation, context length
- Source: Ollama FAQ (KV cache, Flash Attention)
- Source: Lost in the Middle (arXiv 2023)
- Source: RULER, long-context benchmark (arXiv 2024)
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.