Intermediate 11 minConfiguration

Temperature, top-p, top-k: the parameters

Direct response

Top-p (nucleus sampling) keeps only the most likely tokens whose cumulative probability reaches p, for example 0.9. Top-k keeps a fixed number of candidates, for example 40. Temperature controls sampling randomness: near 0, the model almost always chooses the most likely token. Ollama starts with a temperature of 0.8, top-k 40, and top-p 0.9.

Three numbers appear in every local model setting: temperature, top-p, and top-k. This guide explains what each one removes or shifts in the list of possible words, with a numerical example, the default values for Ollama and llama.cpp, recommendations published by model vendors, and how to apply them.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Top-p, top-k, temperature: the definition in one table

For each word it produces, a model calculates a probability for every token in its vocabulary. Sampling parameters then decide which one to draw. Top-p and top-k reduce the list of candidates; temperature changes the gap between likely and unlikely ones. None adds knowledge to the model: they only adjust variety and risk.

What each parameter does
ParameterWhat it affectsLow valueHigh valueDefault Ollama
temperatureGap between likely and unlikely tokensMore deterministic responsesMore variety, more errors0,8
top_pList size, determined by cumulative probabilityShort list, conservative outputBroad list, more diversity0,9
top_kMaximum number of candidatesFew candidatesMany candidates40
min_pThreshold relative to the best probabilityLight filteringStrong filtering0.0 (disabled)
repeat_penaltyPenalty on recently used tokensMore repetitionFewer repetitions, higher risk of oddness1.0 (disabled)

These defaults come from the Ollama Modelfile reference. Many models provide their own values in their configuration file, which then override these defaults—hence the value of reading the model page, as we’ll see below.

#How an LLM chooses the next word

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

For each position, the model produces a probability distribution over its entire vocabulary. Always choosing the most probable token—called greedy decoding—produces flat, repetitive text that can sometimes get stuck in a loop. So we introduce random sampling, constrained by the sampling parameters. A fixed seed makes this sampling reproducible: the Ollama documentation specifies that a given seed produces the same text for the same prompt.

The order in which filters are applied matters. In llama.cpp, the GGUF inference engine also used by LM Studio, the default order is: penalties, DRY, top-n sigma, top-k, typical-p, top-p, min-p, XTC, then temperature last. In other words, top-k and top-p operate on the original probabilities, and temperature is used only to sample from what remains. Many tutorials state the opposite; the actual order may differ from one engine to another, and is configured in llama.cpp with the --samplers parameter.

i
What changes for you
Because temperature is applied after the filters in llama.cpp, increasing the temperature does not bring back absurd tokens that top-k or top-p have already excluded; it redistributes the probabilities among the remaining candidates.

#A numerical example: what each setting retains

Let's take seven candidates for the next word, with probabilities of 0.45; 0.25; 0.12; 0.08; 0.05; 0.03; 0.02. This table is an illustrative calculation, not a measurement on a real model.

Which candidates remain, according to the filter
FilterRuleRetained candidatesCoverage probability
NoneAll the vocabulary7100 %
top_k = 3The 3 best382 %
top_p = 0.9We stop when the cumulative total reaches 0.904 (0,45 + 0,25 + 0,12 + 0,08)90 %
top_p = 0.7Cumulative total: 0.70270 %
min_p = 0.1Keep the tokens within 10% of the best result, meaning a threshold of 0.045595 %

The key takeaway: top-k always keeps the same number of candidates, regardless of the situation; top-p and min-p adapt. When the model is very confident (one token at 0.95), top-p may keep only one or two candidates; when it is choosing among ten words, it keeps more. That is why top-p is often preferred over top-k.

#1. Temperature: the randomness knob

Before sampling, temperature divides the model's logits. At low temperature, the gap between tokens widens: the most likely token overwhelms the others. At high temperature, the gap narrows. At 0, the model always chooses the most likely token and produces identical outputs every time, as the llama.cpp documentation confirms.

Effect of temperature on the previous example (illustrative calculation)
CandidateProbability of originT = 0,5T = 1,5
1er45 %70 %34 %
2e25 %22 %23 %
3e12 %5 %14 %
4e8 %2 %11 %
5e à 7e10% in total1,3 %17,8 %
0 à 0,3
Extraction, classification, faithful summarization: you want the same answer every time.
0,4 à 0,7
Technical writing, translation, support: variety without drift.
0,8
Default value for Ollama and llama.cpp: general conversation.
1 and up
Creativity, brainstorming; greater risk of errors and digressions.
!
Temperature 0: not always the best idea for code
For outputs to be parsed (JSON, classification), a low temperature is reasonable. But developers of reasoning models often recommend higher values: the Qwen3.5 model card gives 0.6 for code in reasoning mode and 1.0 for general tasks. Read the model card before forcing 0.

#2. Top-p: the adaptive filter

Top-p, or nucleus sampling, retains the most probable tokens until their cumulative probability reaches p. With 0.9, it discards the 10% probability tail; with 1.0, it filters nothing. Ollama's documentation notes that a higher value, such as 0.95, produces more varied text, while a lower one, such as 0.5, produces more focused and conservative text.

0,5 à 0,7
Very conservative: safe answers, sometimes monotonous.
0,9
Ollama default: cuts off the tail without stifling variety.
0,95
The default value for llama-server and the value recommended by several model cards.
1,0
Filter disabled: only the temperature remains active.

#3. Top-k: the fixed ceiling

Top-k keeps only the k most likely tokens. Ollama's documentation describes this parameter as a way to reduce the likelihood of producing nonsense: higher (100) means more diverse responses; lower (10) means more conservative ones. Its default is 40, as in llama.cpp, where 0 disables it.

Top-k, top-p, or min-p: which one to use
FilterAdvantageLimitWhen to use it
top_kSimple, bounds computationIgnores the shape of the distribution: too many candidates when the model is confident, not enough when it is uncertainBroad safety margin (20 to 100)
top_pAdapts to the model's confidenceCan let a long tail through if the distribution is flatPrimary diversity setting
min_pThreshold relative to the best probability, stable even at high temperatureLess well known, not enabled by default in OllamaAn alternative to top_p, often with a higher temperature

For the “top p vs. top k” query: use top-p as the primary setting, keep top-k as a safeguard, and don’t adjust both in the same trial, or you won’t know which one caused the difference.

#Defaults and vendor recommendations

An engine's defaults are a generic compromise. Model publishers often provide values tailored to their models, and these may differ from the defaults. Here's what the Qwen3.5 model card recommends for its models, compared with the Ollama defaults (temperature 0.8, top-p 0.9, top-k 40).

Recommended settings in the Qwen3.5-9B model card on Hugging Face
ModeTemperaturetop_ptop_kpresence_penalty
Reasoning, general tasks1,00,95201,5
Reasoning, precise code0,60,95200,0
Without reasoning, general tasks0,70,8201,5

These values apply to this model family, not to all models. Other vendors' model cards provide different values: look for the section on recommended sampling parameters in the Hugging Face model card or on the model's page in the Ollama library. The Qwen model card also notes that support for these parameters varies by inference engine.

#Penalties: repetition, frequency, presence

Local models, especially small or heavily quantized ones, sometimes loop over the same phrases. Penalties are the remedy, but their drawbacks are often poorly understood. In Ollama, repeat_penalty defaults to 1.0, meaning it is disabled; the documentation gives 1.5 as an example of a value that penalizes more strongly. The often-cited value of 1.1 is a choice to make, not a default.

repeat_penalty
Penalizes recently seen tokens; repeat_last_n (64 by default) sets the window examined.
frequency_penalty
Penalty proportional to the number of occurrences, for long generations.
presence_penalty
A penalty as soon as a token has appeared, to encourage a more varied vocabulary; Qwen recommends 1.5 in several of its modes, with the risk of language mixing if the value is too high.
→
Increase penalties in small increments
A strong penalty pushes the model toward rare or unexpected words. Increase repeat_penalty in increments of 0.05 to 0.1 and stop as soon as the loops disappear.

#Apply these settings in Ollama

There are two methods: a Modelfile, which sets values for a named model, or the options field of an API request, which accepts the same parameters as the Modelfile. The Modelfile is suited to daily use; the API is suited to a program.

Modelfile: a custom model
FROM llama3.2
PARAMETER temperature 0.3
PARAMETER top_p 0.9
PARAMETER top_k 40
Create and then run this model
ollama create llama-precis -f ./Modelfile
ollama run llama-precis
Same setting in an API request
curl http://localhost:11434/api/generate -d '{"model":"llama3.2","prompt":"Résume en une phrase : ...","stream":false,"options":{"temperature":0.3,"top_p":0.9}}'

For more on Modelfiles (system prompt, context, template), see the customization guide. For LM Studio and Jan, the same settings are configured in the model settings panel.

#Reasoning models: don't solve everything with temperature

Reasoning models first produce a chain of thought before their answer. On Ollama, these models expose a separate reasoning field, and the available commands vary by model: the documentation recommends querying the show API to find the accepted values and the default. For gpt-oss, for example, the levels are low, medium, and high, with medium as the default.

To reduce unnecessary overthinking on a simple question, adjust this reasoning level first, when available, rather than the temperature. Temperature should follow the model's specifications: reasoning that is too cold can freeze up, while reasoning that is too hot can wander.

#Starting presets by use case

Starting points to adjust, excluding models with specific recommendations
UsageTemperaturetop_pOther settings
Extraction, classification, JSON0 à 0,20,9Strict format requested in the prompt
Faithful summary0,2 à 0,30,9repeat_penalty slightly above 1
General conversation0,7 à 0,80,9Engine limitations
Writing, ideas0,8 à 1,00,95moderate presence_penalty
Fiction, creativity1,00,95Control digressions

These presets are starting points, not measurements. Change one parameter at a time, run the same prompt three times, and compare: if three attempts produce the same answer when you expect variety, the temperature is too low or the seed is fixed.

#Symptoms and settings to try

The model repeats the same sentence
Set repeat_penalty to around 1.1 or use presence_penalty, or shorten generation with num_predict.
Flat, generic responses
Increase the temperature by one step (0.7 to 0.9) or loosen top_p.
The model makes up facts
Lower the temperature, but be aware that this isn't enough: see the guide to hallucinations.
Broken JSON
Low temperature and an explicitly requested format; use Ollama's structured output mode when available.
Mixed languages or strange words
Lower presence_penalty or repeat_penalty, or slightly increase the model’s quantization.
Frequently asked questions about top-p, top-k, and temperature
What is top-p in an LLM?+
Top-p is a filter that keeps only the most probable tokens whose cumulative probability reaches the p threshold, such as 0.9. The rest are discarded before sampling. Unlike top-k, the number of candidates varies: fewer when the model is confident, more when it is uncertain.
What's the difference between top-p and top-k?+
Top-k always keeps a fixed number of candidates, such as 40, regardless of the situation. Top-p keeps a variable number of candidates based on cumulative probability. Top-p therefore adapts better to the model’s confidence; top-k serves as a simple safeguard against highly improbable tokens.
Should you tune temperature and top-p at the same time?+
It is better to change one parameter at a time to understand its effect. In practice, leave top-p and top-k at their default values or those listed in the model card, and adjust the temperature. If you change several settings at once, you will not know which one caused the difference.
What temperature should you use for code?+
For a model without reasoning, a low value (0 to 0.2) produces reproducible outputs. For a reasoning model, follow the publisher's documentation: Qwen3.5's recommends 0.6 for precise code in reasoning mode. Test on your own use cases and verify that the code compiles or passes your tests.
What are the default values in Ollama?+
The Modelfile documentation specifies a temperature of 0.8, a top-k of 40, a top-p of 0.9, a min-p of 0.0 (disabled), and a repetition penalty of 1.0 (disabled). A model may include its own values and override these; ollama show --modelfile lets you view them.
What is min-p, and should you use it?+
Min-p eliminates tokens whose probability is lower than a fraction of the best candidate's probability: with 0.05 and a best token at 0.9, the threshold is 0.045. It is an alternative to top-p, disabled by default in Ollama. Use it only if the model card recommends it.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.