Temperature, top-p, top-k: the parameters
Top-p (nucleus sampling) keeps only the most likely tokens whose cumulative probability reaches p, for example 0.9. Top-k keeps a fixed number of candidates, for example 40. Temperature controls sampling randomness: near 0, the model almost always chooses the most likely token. Ollama starts with a temperature of 0.8, top-k 40, and top-p 0.9.
Three numbers appear in every local model setting: temperature, top-p, and top-k. This guide explains what each one removes or shifts in the list of possible words, with a numerical example, the default values for Ollama and llama.cpp, recommendations published by model vendors, and how to apply them.
#Top-p, top-k, temperature: the definition in one table
For each word it produces, a model calculates a probability for every token in its vocabulary. Sampling parameters then decide which one to draw. Top-p and top-k reduce the list of candidates; temperature changes the gap between likely and unlikely ones. None adds knowledge to the model: they only adjust variety and risk.
| Parameter | What it affects | Low value | High value | Default Ollama |
|---|---|---|---|---|
| temperature | Gap between likely and unlikely tokens | More deterministic responses | More variety, more errors | 0,8 |
| top_p | List size, determined by cumulative probability | Short list, conservative output | Broad list, more diversity | 0,9 |
| top_k | Maximum number of candidates | Few candidates | Many candidates | 40 |
| min_p | Threshold relative to the best probability | Light filtering | Strong filtering | 0.0 (disabled) |
| repeat_penalty | Penalty on recently used tokens | More repetition | Fewer repetitions, higher risk of oddness | 1.0 (disabled) |
These defaults come from the Ollama Modelfile reference. Many models provide their own values in their configuration file, which then override these defaults—hence the value of reading the model page, as we’ll see below.
#How an LLM chooses the next word
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
For each position, the model produces a probability distribution over its entire vocabulary. Always choosing the most probable token—called greedy decoding—produces flat, repetitive text that can sometimes get stuck in a loop. So we introduce random sampling, constrained by the sampling parameters. A fixed seed makes this sampling reproducible: the Ollama documentation specifies that a given seed produces the same text for the same prompt.
The order in which filters are applied matters. In llama.cpp, the GGUF inference engine also used by LM Studio, the default order is: penalties, DRY, top-n sigma, top-k, typical-p, top-p, min-p, XTC, then temperature last. In other words, top-k and top-p operate on the original probabilities, and temperature is used only to sample from what remains. Many tutorials state the opposite; the actual order may differ from one engine to another, and is configured in llama.cpp with the --samplers parameter.
#A numerical example: what each setting retains
Let's take seven candidates for the next word, with probabilities of 0.45; 0.25; 0.12; 0.08; 0.05; 0.03; 0.02. This table is an illustrative calculation, not a measurement on a real model.
| Filter | Rule | Retained candidates | Coverage probability |
|---|---|---|---|
| None | All the vocabulary | 7 | 100 % |
| top_k = 3 | The 3 best | 3 | 82 % |
| top_p = 0.9 | We stop when the cumulative total reaches 0.90 | 4 (0,45 + 0,25 + 0,12 + 0,08) | 90 % |
| top_p = 0.7 | Cumulative total: 0.70 | 2 | 70 % |
| min_p = 0.1 | Keep the tokens within 10% of the best result, meaning a threshold of 0.045 | 5 | 95 % |
The key takeaway: top-k always keeps the same number of candidates, regardless of the situation; top-p and min-p adapt. When the model is very confident (one token at 0.95), top-p may keep only one or two candidates; when it is choosing among ten words, it keeps more. That is why top-p is often preferred over top-k.
#1. Temperature: the randomness knob
Before sampling, temperature divides the model's logits. At low temperature, the gap between tokens widens: the most likely token overwhelms the others. At high temperature, the gap narrows. At 0, the model always chooses the most likely token and produces identical outputs every time, as the llama.cpp documentation confirms.
| Candidate | Probability of origin | T = 0,5 | T = 1,5 |
|---|---|---|---|
| 1er | 45 % | 70 % | 34 % |
| 2e | 25 % | 22 % | 23 % |
| 3e | 12 % | 5 % | 14 % |
| 4e | 8 % | 2 % | 11 % |
| 5e à 7e | 10% in total | 1,3 % | 17,8 % |
- 0 à 0,3
- Extraction, classification, faithful summarization: you want the same answer every time.
- 0,4 à 0,7
- Technical writing, translation, support: variety without drift.
- 0,8
- Default value for Ollama and llama.cpp: general conversation.
- 1 and up
- Creativity, brainstorming; greater risk of errors and digressions.
#2. Top-p: the adaptive filter
Top-p, or nucleus sampling, retains the most probable tokens until their cumulative probability reaches p. With 0.9, it discards the 10% probability tail; with 1.0, it filters nothing. Ollama's documentation notes that a higher value, such as 0.95, produces more varied text, while a lower one, such as 0.5, produces more focused and conservative text.
- 0,5 à 0,7
- Very conservative: safe answers, sometimes monotonous.
- 0,9
- Ollama default: cuts off the tail without stifling variety.
- 0,95
- The default value for llama-server and the value recommended by several model cards.
- 1,0
- Filter disabled: only the temperature remains active.
#3. Top-k: the fixed ceiling
Top-k keeps only the k most likely tokens. Ollama's documentation describes this parameter as a way to reduce the likelihood of producing nonsense: higher (100) means more diverse responses; lower (10) means more conservative ones. Its default is 40, as in llama.cpp, where 0 disables it.
| Filter | Advantage | Limit | When to use it |
|---|---|---|---|
| top_k | Simple, bounds computation | Ignores the shape of the distribution: too many candidates when the model is confident, not enough when it is uncertain | Broad safety margin (20 to 100) |
| top_p | Adapts to the model's confidence | Can let a long tail through if the distribution is flat | Primary diversity setting |
| min_p | Threshold relative to the best probability, stable even at high temperature | Less well known, not enabled by default in Ollama | An alternative to top_p, often with a higher temperature |
For the “top p vs. top k” query: use top-p as the primary setting, keep top-k as a safeguard, and don’t adjust both in the same trial, or you won’t know which one caused the difference.
#Defaults and vendor recommendations
An engine's defaults are a generic compromise. Model publishers often provide values tailored to their models, and these may differ from the defaults. Here's what the Qwen3.5 model card recommends for its models, compared with the Ollama defaults (temperature 0.8, top-p 0.9, top-k 40).
| Mode | Temperature | top_p | top_k | presence_penalty |
|---|---|---|---|---|
| Reasoning, general tasks | 1,0 | 0,95 | 20 | 1,5 |
| Reasoning, precise code | 0,6 | 0,95 | 20 | 0,0 |
| Without reasoning, general tasks | 0,7 | 0,8 | 20 | 1,5 |
These values apply to this model family, not to all models. Other vendors' model cards provide different values: look for the section on recommended sampling parameters in the Hugging Face model card or on the model's page in the Ollama library. The Qwen model card also notes that support for these parameters varies by inference engine.
#Penalties: repetition, frequency, presence
Local models, especially small or heavily quantized ones, sometimes loop over the same phrases. Penalties are the remedy, but their drawbacks are often poorly understood. In Ollama, repeat_penalty defaults to 1.0, meaning it is disabled; the documentation gives 1.5 as an example of a value that penalizes more strongly. The often-cited value of 1.1 is a choice to make, not a default.
- repeat_penalty
- Penalizes recently seen tokens; repeat_last_n (64 by default) sets the window examined.
- frequency_penalty
- Penalty proportional to the number of occurrences, for long generations.
- presence_penalty
- A penalty as soon as a token has appeared, to encourage a more varied vocabulary; Qwen recommends 1.5 in several of its modes, with the risk of language mixing if the value is too high.
#Apply these settings in Ollama
There are two methods: a Modelfile, which sets values for a named model, or the options field of an API request, which accepts the same parameters as the Modelfile. The Modelfile is suited to daily use; the API is suited to a program.
For more on Modelfiles (system prompt, context, template), see the customization guide. For LM Studio and Jan, the same settings are configured in the model settings panel.
#Reasoning models: don't solve everything with temperature
Reasoning models first produce a chain of thought before their answer. On Ollama, these models expose a separate reasoning field, and the available commands vary by model: the documentation recommends querying the show API to find the accepted values and the default. For gpt-oss, for example, the levels are low, medium, and high, with medium as the default.
To reduce unnecessary overthinking on a simple question, adjust this reasoning level first, when available, rather than the temperature. Temperature should follow the model's specifications: reasoning that is too cold can freeze up, while reasoning that is too hot can wander.
#Starting presets by use case
| Usage | Temperature | top_p | Other settings |
|---|---|---|---|
| Extraction, classification, JSON | 0 à 0,2 | 0,9 | Strict format requested in the prompt |
| Faithful summary | 0,2 à 0,3 | 0,9 | repeat_penalty slightly above 1 |
| General conversation | 0,7 à 0,8 | 0,9 | Engine limitations |
| Writing, ideas | 0,8 à 1,0 | 0,95 | moderate presence_penalty |
| Fiction, creativity | 1,0 | 0,95 | Control digressions |
These presets are starting points, not measurements. Change one parameter at a time, run the same prompt three times, and compare: if three attempts produce the same answer when you expect variety, the temperature is too low or the seed is fixed.
#Symptoms and settings to try
- The model repeats the same sentence
- Set repeat_penalty to around 1.1 or use presence_penalty, or shorten generation with num_predict.
- Flat, generic responses
- Increase the temperature by one step (0.7 to 0.9) or loosen top_p.
- The model makes up facts
- Lower the temperature, but be aware that this isn't enough: see the guide to hallucinations.
- Broken JSON
- Low temperature and an explicitly requested format; use Ollama's structured output mode when available.
- Mixed languages or strange words
- Lower presence_penalty or repeat_penalty, or slightly increase the model’s quantization.
What is top-p in an LLM?+
What's the difference between top-p and top-k?+
Should you tune temperature and top-p at the same time?+
What temperature should you use for code?+
What are the default values in Ollama?+
What is min-p, and should you use it?+
- Ollama Modelfile: customize a model
- Prompting basics
- Master system prompts
- Understanding the context window
- Limit hallucinations in a local LLM
- Source: Ollama Modelfile reference
- Source: Qwen3.5-9B specification
- Source: llama.cpp, llama-server options
- Source: Ollama, reasoning models
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.