Hallucinations: why your local LLM makes things up and how limiter
An AI hallucination is a false answer stated with the confidence of a correct one: an invented date, a nonexistent citation, an imaginary Python function. On a local LLM, the phenomenon is the same as with large cloud models, sometimes more pronounced with smaller quantizations. This guide explains where these inventions come from, without jargon, then lists the concrete settings, prompts, and safeguards that can reduce them—and especially the cases where you should never trust the machine.
#What is a hallucination
The term “hallucination” refers to any statement produced by the model that is false, fabricated, or impossible to verify, but presented as a fact. This isn’t a bug in the computer-science sense: the program works perfectly and generates plausible text. The problem is that “plausible” and “true” aren’t the same thing.
In practice, a hallucination takes several forms: a precise but false figure (“the population of this city is 47,312”), a nonexistent source (“according to the study by Dupont et al., 2019”), an imaginary API or command (`ollama sync --cloud`), or a mixture of real elements recombined incorrectly. The common thread: the tone is always confident.
#Why the model makes things up
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
An LLM is trained for a single task: completing text in a statistically plausible way. It has absorbed billions of sentences and extracted patterns from them. When you ask a question, it does not “consult” a knowledge base: it generates, word by word, the most likely continuation given your query and its training.
- No reliable factual memory
- Knowledge is diluted across the network's weights, not stored like it is in a database. The model “remembers” a trend, not an exact fact. Precise details (dates, figures, rare proper names) are the first to become distorted.
- Training frozen in time
- The model knows nothing after its training cutoff date. When asked about a recent event, it does not say “I don’t know”: it extrapolates from what it knows, producing fabrications.
- Response bias
- An LLM is optimized to answer, not to stay silent. When faced with a question it does not know the answer to, the most likely continuation is often a confidently worded response rather than an admission of ignorance.
- The effect of quantization
- Compressing a model to Q4_K_M saves VRAM but slightly reduces accuracy. On highly specific factual tasks, aggressive quantization can increase the hallucination rate compared with Q8_0 or FP16.
- The randomness of sampling
- For each word, the model randomly selects from the likely candidates. The more “creative” this selection is (higher temperature), the more it can drift toward unlikely—and therefore false—continuations.
In other words, hallucination is not an accident: it is the normal behavior of a system that produces plausible text without access to the truth. You do not eliminate it; you reduce and control it.
#Spot an invention
Before fixing something, you need to know how to detect it. Some signals should immediately put you on alert.
- A suspicious level of precision
- Figures accurate to the unit, exact dates, precise percentages on a niche topic: the more precise something is without a source, the more questionable it is.
- Citations and links
- Article titles, author names, URLs, page numbers, and legal references. LLMs invent sources with perfect credibility. No reference produced by a model should be considered real without verification.
- Code that “should” exist
- Function or command-line option names that seem logical but don't exist. The model completes them by analogy with neighboring APIs.
- A response that changes
- Ask exactly the same question in a new conversation. If the factual answer varies from one attempt to the next, the model is guessing rather than knowing.
#Settings that limit hallucinations
Generation parameters control the trade-off between creativity and reliability. For anything involving facts, code, or information extraction, you want deterministic, cautious output.
- temperature
- The main lever. At 0, the model always chooses the most likely word: stable, conservative responses. Increase it to 0.7–1.0 for creative writing, but lower it to 0.1–0.3 for factual tasks.
- top_p
- Limits sampling to candidates whose cumulative probability reaches a given value. A low value (0.1–0.5) cuts off the long tail of unlikely words, which are often responsible for derailments.
- top_k
- Limits the number of candidates considered at each step. A low top_k (10–20) reduces the window in which the model can go off the rails.
- seed
- Fixing a seed makes generation reproducible: with identical settings, you get the same output. Essential for testing whether a response is stable or random.
- num_ctx
- Context size. If it's too short, the model “forgets” the beginning and may incorrectly recombine fragments. Make sure your reference documents fit within the window.
With Ollama, these settings can be passed on the fly in the API call or fixed in a Modelfile. Here's an intentionally cautious, factual example of a call to the local daemon:
To make these settings permanent for a model, add them to a Modelfile and create a variant dedicated to factual tasks:
#Safety prompts
How you phrase your request greatly affects the rate of hallucinations. The idea: explicitly allow uncertainty and prohibit invention.
- Allowing yourself not to know
- “If you're not sure, answer: I don't know.” Without that permission, the model will fill the gap with an invention.
- Require grounding
- “Answer only from the text below. Do not use any outside knowledge.” This forces the model to limit itself to the provided context.
- Request sources in the text
- “For every claim, cite the exact sentence from the document that supports it.” If the model cannot find supporting evidence, the absence becomes visible.
- Separate facts from assumptions
- “Distinguish what is established from what is your own assumption.” This pushes the model to label its own uncertainties.
- Break down complex tasks
- A question broken into several explicit steps leaves less room for improvisation than one large open-ended question.
#Ground responses in your own documents
The most effective technique against AI hallucinations is still RAG (Retrieval-Augmented Generation): instead of letting the model draw from its diffuse memory, you provide it with relevant passages from your documents and ask it to answer based on those passages only. The model shifts from being a “source” to being a “reader.”
- 01Index your documentsYour files (PDFs, notes, internal docs) are split into chunks, converted into vectors by an embeddings model, and stored in a vector database.
- 02Find the relevant passagesFor each question, the system searches for and retrieves the chunks that are semantically closest to the query.
- 03Inject into the contextThe retrieved passages are inserted into the prompt, along with a strict instruction to answer only from these excerpts.
- 04Generate with citationsThe model drafts the answer based on the provided excerpts, ideally citing which ones. If they do not contain the answer, it must say so.
Open WebUI, connected to the Ollama daemon (http://localhost:11434), offers built-in RAG: you upload documents and reference them with `#` in the conversation. For custom needs, a vector database and a homegrown pipeline offer more control.
#Always verify
No technique makes an LLM 100% reliable. Verification is therefore not optional; it is a workflow step. It should be proportionate to the stakes.
- Cross-check against a real source
- Any factual information intended for use (number, date, quotation) must be verified against a primary source. The model is a starting point, never a reference.
- Test the code before trusting it
- Run what the model produces. A function that does not exist will raise an immediate error. Never copy critical code without running it first.
- Compare two generations
- Ask the question again with a different seed or in a new session. Stable points are more reliable; points that vary are suspect.
- Use a second model
- Having one model review another model’s response (or that of a larger variant) brings obvious inconsistencies to light.
- Keep a human in the loop
- For any decision with consequences (health, legal, financial, or security-related), final validation remains human. No exceptions.
#Cases where you should never trust it
Some domains concentrate the most dangerous hallucinations. In those areas, treat all model output as an unverified draft, assumed false by default until proven otherwise.
- Health and medications
- Dosages, interactions, diagnoses. An invention can be dangerous. The model is not a healthcare professional.
- Law and tax
- Statutes, case law, filing requirements. LLMs invent legal references with deceptively realistic details.
- Precise figures and statistics
- Populations, rates, amounts, exact dates. Numerical details are a structural weakness of models.
- Recent events
- Everything after the training cutoff. The model will fill it in by extrapolation without telling you.
- Citations and references
- Titles, authors, URLs, page numbers. Always verify them systematically, without exception, before reusing anything.
- Niche people and facts
- Biographies of little-known people, obscure details: the model recombines fragments and makes up the rest.
#Go further
Reducing hallucinations is mainly about combining the right settings with the right grounding. These guides complement the approach:
- Temperature, top-p, top-k: the parameters
- To gain detailed control over the sampling levers discussed here and fine-tune the creativity/reliability tradeoff.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To understand the effect of compression on factual accuracy and decide when reliability takes priority.
- Master system prompts
- For a deeper look at caution prompts and locking in anti-hallucination behavior by default.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.