Beginner 9 minConcepts

Hallucinations: why your local LLM makes things up and how limiter

An AI hallucination is a false answer stated with the confidence of a correct one: an invented date, a nonexistent citation, an imaginary Python function. On a local LLM, the phenomenon is the same as with large cloud models, sometimes more pronounced with smaller quantizations. This guide explains where these inventions come from, without jargon, then lists the concrete settings, prompts, and safeguards that can reduce them—and especially the cases where you should never trust the machine.

By Clara M.·Update 2026-08-03·Tested on Windows, macOS, and Linux

#What is a hallucination

The term “hallucination” refers to any statement produced by the model that is false, fabricated, or impossible to verify, but presented as a fact. This isn’t a bug in the computer-science sense: the program works perfectly and generates plausible text. The problem is that “plausible” and “true” aren’t the same thing.

In practice, a hallucination takes several forms: a precise but false figure (“the population of this city is 47,312”), a nonexistent source (“according to the study by Dupont et al., 2019”), an imaginary API or command (`ollama sync --cloud`), or a mixture of real elements recombined incorrectly. The common thread: the tone is always confident.

i
This isn't a lie
The model has no intention to deceive. It doesn't even have a concept of true or false. It predicts the most likely next word. When truth and probability diverge, probability wins.

#Why the model makes things up

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

An LLM is trained for a single task: completing text in a statistically plausible way. It has absorbed billions of sentences and extracted patterns from them. When you ask a question, it does not “consult” a knowledge base: it generates, word by word, the most likely continuation given your query and its training.

No reliable factual memory
Knowledge is diluted across the network's weights, not stored like it is in a database. The model “remembers” a trend, not an exact fact. Precise details (dates, figures, rare proper names) are the first to become distorted.
Training frozen in time
The model knows nothing after its training cutoff date. When asked about a recent event, it does not say “I don’t know”: it extrapolates from what it knows, producing fabrications.
Response bias
An LLM is optimized to answer, not to stay silent. When faced with a question it does not know the answer to, the most likely continuation is often a confidently worded response rather than an admission of ignorance.
The effect of quantization
Compressing a model to Q4_K_M saves VRAM but slightly reduces accuracy. On highly specific factual tasks, aggressive quantization can increase the hallucination rate compared with Q8_0 or FP16.
The randomness of sampling
For each word, the model randomly selects from the likely candidates. The more “creative” this selection is (higher temperature), the more it can drift toward unlikely—and therefore false—continuations.

In other words, hallucination is not an accident: it is the normal behavior of a system that produces plausible text without access to the truth. You do not eliminate it; you reduce and control it.

#Spot an invention

Before fixing something, you need to know how to detect it. Some signals should immediately put you on alert.

A suspicious level of precision
Figures accurate to the unit, exact dates, precise percentages on a niche topic: the more precise something is without a source, the more questionable it is.
Citations and links
Article titles, author names, URLs, page numbers, and legal references. LLMs invent sources with perfect credibility. No reference produced by a model should be considered real without verification.
Code that “should” exist
Function or command-line option names that seem logical but don't exist. The model completes them by analogy with neighboring APIs.
A response that changes
Ask exactly the same question in a new conversation. If the factual answer varies from one attempt to the next, the model is guessing rather than knowing.
→
The rewriting test
Ask the model, « are you sure? cite your exact source » or ask the question another way. A solid fact holds up; an invention often contradicts itself or collapses into « I can't confirm ».

#Settings that limit hallucinations

Generation parameters control the trade-off between creativity and reliability. For anything involving facts, code, or information extraction, you want deterministic, cautious output.

temperature
The main lever. At 0, the model always chooses the most likely word: stable, conservative responses. Increase it to 0.7–1.0 for creative writing, but lower it to 0.1–0.3 for factual tasks.
top_p
Limits sampling to candidates whose cumulative probability reaches a given value. A low value (0.1–0.5) cuts off the long tail of unlikely words, which are often responsible for derailments.
top_k
Limits the number of candidates considered at each step. A low top_k (10–20) reduces the window in which the model can go off the rails.
seed
Fixing a seed makes generation reproducible: with identical settings, you get the same output. Essential for testing whether a response is stable or random.
num_ctx
Context size. If it's too short, the model “forgets” the beginning and may incorrectly recombine fragments. Make sure your reference documents fit within the window.

With Ollama, these settings can be passed on the fly in the API call or fixed in a Modelfile. Here's an intentionally cautious, factual example of a call to the local daemon:

Terminal
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5:9b",
  "prompt": "Résume ce texte sans rien ajouter : ...",
  "stream": false,
  "options": {
    "temperature": 0.2,
    "top_p": 0.5,
    "top_k": 20,
    "seed": 42
  }
}'

To make these settings permanent for a model, add them to a Modelfile and create a variant dedicated to factual tasks:

Modelfile
FROM mistral
PARAMETER temperature 0.2
PARAMETER top_p 0.5
PARAMETER top_k 20
SYSTEM "Tu réponds uniquement à partir des informations fournies. Si tu ne sais pas, tu dis 'Je ne sais pas'. Tu n'inventes jamais de source, de chiffre ni de citation."
Terminal
ollama create mistral-prudent -f ./Modelfile
ollama run mistral-prudent
!
Low temperature ≠ guaranteed truth
Lowering the temperature makes the response stable and conservative, not accurate. A model can hallucinate in a perfectly deterministic way at temperature 0. Settings reduce noise; they don't create knowledge that doesn't exist.

#Safety prompts

How you phrase your request greatly affects the rate of hallucinations. The idea: explicitly allow uncertainty and prohibit invention.

Allowing yourself not to know
“If you're not sure, answer: I don't know.” Without that permission, the model will fill the gap with an invention.
Require grounding
“Answer only from the text below. Do not use any outside knowledge.” This forces the model to limit itself to the provided context.
Request sources in the text
“For every claim, cite the exact sentence from the document that supports it.” If the model cannot find supporting evidence, the absence becomes visible.
Separate facts from assumptions
“Distinguish what is established from what is your own assumption.” This pushes the model to label its own uncertainties.
Break down complex tasks
A question broken into several explicit steps leaves less room for improvisation than one large open-ended question.
Anti-hallucination system prompt
Tu es un assistant factuel et prudent.
Règles :
1. Réponds uniquement à partir du contexte fourni par l'utilisateur.
2. Si l'information n'y figure pas, réponds exactement : « Je ne sais pas ».
3. N'invente jamais de chiffre, de date, de nom, de source ni d'URL.
4. Pour chaque affirmation, cite la phrase du contexte qui la justifie.
5. Distingue clairement les faits des hypothèses.
→
The simple “I don't know” changes everything
Adding an instruction that allows the model to admit ignorance is the most cost-effective step. Many fabrications happen solely because the model thinks it must answer at all costs.

#Ground responses in your own documents

The most effective technique against AI hallucinations is still RAG (Retrieval-Augmented Generation): instead of letting the model draw from its diffuse memory, you provide it with relevant passages from your documents and ask it to answer based on those passages only. The model shifts from being a “source” to being a “reader.”

  1. 01
    Index your documents
    Your files (PDFs, notes, internal docs) are split into chunks, converted into vectors by an embeddings model, and stored in a vector database.
  2. 02
    Find the relevant passages
    For each question, the system searches for and retrieves the chunks that are semantically closest to the query.
  3. 03
    Inject into the context
    The retrieved passages are inserted into the prompt, along with a strict instruction to answer only from these excerpts.
  4. 04
    Generate with citations
    The model drafts the answer based on the provided excerpts, ideally citing which ones. If they do not contain the answer, it must say so.

Open WebUI, connected to the Ollama daemon (http://localhost:11434), offers built-in RAG: you upload documents and reference them with `#` in the conversation. For custom needs, a vector database and a homegrown pipeline offer more control.

i
RAG reduces it; it does not eliminate it
Even when grounded, a model can misinterpret a passage or extrapolate beyond it. Chunking and retrieval quality matter just as much as the model. A poorly tuned RAG that retrieves the wrong excerpts produces confident… and false answers.

#Always verify

No technique makes an LLM 100% reliable. Verification is therefore not optional; it is a workflow step. It should be proportionate to the stakes.

Cross-check against a real source
Any factual information intended for use (number, date, quotation) must be verified against a primary source. The model is a starting point, never a reference.
Test the code before trusting it
Run what the model produces. A function that does not exist will raise an immediate error. Never copy critical code without running it first.
Compare two generations
Ask the question again with a different seed or in a new session. Stable points are more reliable; points that vary are suspect.
Use a second model
Having one model review another model’s response (or that of a larger variant) brings obvious inconsistencies to light.
Keep a human in the loop
For any decision with consequences (health, legal, financial, or security-related), final validation remains human. No exceptions.

#Cases where you should never trust it

Some domains concentrate the most dangerous hallucinations. In those areas, treat all model output as an unverified draft, assumed false by default until proven otherwise.

Health and medications
Dosages, interactions, diagnoses. An invention can be dangerous. The model is not a healthcare professional.
Law and tax
Statutes, case law, filing requirements. LLMs invent legal references with deceptively realistic details.
Precise figures and statistics
Populations, rates, amounts, exact dates. Numerical details are a structural weakness of models.
Recent events
Everything after the training cutoff. The model will fill it in by extrapolation without telling you.
Citations and references
Titles, authors, URLs, page numbers. Always verify them systematically, without exception, before reusing anything.
Niche people and facts
Biographies of little-known people, obscure details: the model recombines fragments and makes up the rest.
!
The golden rule
A local LLM is an excellent writing, brainstorming, and rephrasing assistant. It is not a verified knowledge base. The higher the stakes, the stricter human verification must be.

#Go further

Reducing hallucinations is mainly about combining the right settings with the right grounding. These guides complement the approach:

Temperature, top-p, top-k: the parameters
To gain detailed control over the sampling levers discussed here and fine-tune the creativity/reliability tradeoff.
Choose your quantization (Q4, Q5, Q8, FP16)
To understand the effect of compression on factual accuracy and decide when reliability takes priority.
Master system prompts
For a deeper look at caution prompts and locking in anti-hallucination behavior by default.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.