The basics of prompting
To query an LLM, write a complete instruction as if you were briefing a colleague who knows neither your project nor your habits: the precise task, the relevant context, the data to process clearly separated, and the expected response format. You can send it through a chat interface, the command line, or the API. The model cannot infer your intent or the format: anything you do not write will be improvised.
This guide first answers the question in the title, then details the structure of a prompt, the errors that cost the most time, the techniques that work on almost every model, and what changes with a small local model. The principles are based on Anthropic's official prompt engineering documentation and that of Ollama.
#How to query an LLM: three ways, one rule
A language model continues text. What it produces therefore depends almost entirely on what you give it. Anthropic's documentation states a useful golden rule for any model: show your prompt to a colleague who does not know the task and ask them to execute it; if they are lost, the model will be too. It recommends explicitly asking for what you want, including the level of effort, rather than expecting the model to infer it from a vague instruction. This rule applies just as much to an 8-billion-parameter model on your computer as to a cutting-edge model. Parameters, quantization, and the chosen model matter, but a vague prompt ruins any model.
#Three ways to send a question to a local model
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Chat interface
- Open WebUI, LM Studio, Jan: you type, the model responds, and the history is retained. The simplest way to get started.
- Command line
- ollama run suivi du nom du modèle ouvre une session interactive dans le terminal. Pratique pour des essais rapides.
- API
- A program sends an HTTP request to the local server and receives the response as JSON. This is the path for scripts and automations.
The system role message sets the permanent framework, while the user role message contains the request. The guide to system prompts explains what to put in them. Whichever route you choose, the quality of the response depends on what the message contains.
#Anatomy of a good prompt: role, task, context, format
Many guides focus on four elements: role, task, context, and format. This isn't a standard, but a mnemonic that aligns with the documentation's recommendations: be precise, provide context, structure the content, set the format, and, when needed, specify a role. Anthropic's documentation notes that a one-sentence role in the system message already steers tone and behavior.
- Role
- A sentence that situates the model: “You are a rigorous French proofreader.” It sets the register, not the intelligence.
- Task
- The action verb, as precise as possible: “Summarize in five factual bullet points, without adjectives” is better than “Summarize.”
- Context
- The data to process and the reason for the request. Explaining why a constraint exists helps the model generalize it correctly.
- Format
- Length, structure, language, tone. Without these details, the model chooses for itself.
The tags enclosing the contract in the example above are not decorative: the documentation states that XML tags help the model unambiguously distinguish instructions, context, examples, and variable inputs. Choose consistent, descriptive names, and use the same delimiter from one prompt to the next.
#The errors that waste the most time
| Error | Symptom | Correction |
|---|---|---|
| Keyword prompt (“contract summary risks”) | The model does not distinguish between a question and a command | Write a complete sentence with a verb |
| Data mixed with instructions | The model treats part of your instructions as content | Wrap the data in tags |
| Implicit expectations | Unexpected length, language, or tone | Specify the length, language, and tone |
| Prohibitions only (“don't use Markdown”) | The model does it anyway | State what you want: “respond in written paragraphs” |
| Long prompt without structure | Ignored instructions | Number the steps, separate the blocks |
| Constraint without a reason | Awkward application | Explain why: “the text will be read aloud” |
On the penultimate line, Anthropic’s documentation recommends telling the model what to do rather than what not to do, and gives the example of replacing “don’t use Markdown” with a description of the desired format: paragraphs of flowing prose. On the last line, it recommends giving the reason for an instruction: the model derives a general rule from it rather than a literal prohibition.
#Techniques that work on almost all models
#Give examples
According to Anthropic’s documentation, examples are one of the most reliable ways to guide format, tone, and structure. Three to five examples generally produce the best results, provided they are relevant, varied (including edge cases), and enclosed in tags so they are not mistaken for instructions. This is particularly effective for classification, extraction, or reformatting.
#Ask it to reason
For a logic problem, calculation, or code, asking the model to lay out the steps before concluding often improves the result, at the cost of a longer response. Some local models do this on their own: Ollama’s documentation explains that reasoning models return a thinking field separate from the final answer, which you can read, display, or hide. For other models, one sentence is enough: “Explain your reasoning step by step, then give the final answer.”
#Place a long document
Anthropic's documentation recommends that, for inputs longer than 20,000 tokens, you place long documents at the top of the prompt, before the question, and write the question at the end: according to its tests, this can improve response quality by up to 30%, especially for inputs containing multiple documents. This result comes from tests on Claude models: verify it on your local model with your own texts.
#Have the response checked
Ask the model to cite, for every claim, the passage from the provided text that supports it, or to write “unverified” if it cannot find one. This instruction does not eliminate hallucinations, but it makes errors easier to spot. The guide on hallucinations explains its limitations in detail.
#A concrete example: from a vague instruction to a usable prompt
Take a typical request: you paste in meeting minutes and write “summarize this.” The model does not know who the summary is for, how many lines you expect, whether it should list decisions or tasks, or which language to use. It produces a generic text of arbitrary length, which you will correct manually.
| Item | Blurry version | Specified version |
|---|---|---|
| Task | Give me a summary | Summarize this report in five bullet points |
| Public | Not specified | For a director who was not present |
| Expected content | Not specified | Decisions made, followed by actions with an owner and date |
| Format | Free | Bulleted list, one sentence per bullet, 20 words maximum |
| Garde-fou | None | If a date or owner is missing, write “not specified” instead of making it up |
The last line is the most useful with a local model: giving the model an explicit fallback when information is missing reduces hallucinations. Each addition takes only one sentence, and together they turn a result that needs reworking into a usable one. Test it again on three different reports before locking in your prompt: a prompt that succeeds on only one example is fragile. Then keep your best version in a text file: reused as-is, it will save you more time than any wording trick, and it will let you compare two models with the same instructions.
#Enforce an output format
When the response must be read by a program, a simple phrase such as “respond in JSON” is not always enough. Ollama offers structured outputs: you pass a JSON schema in the request’s format field, and the response is constrained to follow it. The documentation recommends also providing the schema in the prompt to anchor the response, lowering the temperature (for example, to 0) for more deterministic results, and defining the schema with Pydantic or Zod so it can be reused for validation. Structured outputs also work through the OpenAI-compatible API with response_format. They are not available for Ollama’s cloud models.
#Iterate without starting from scratch
An imperfect answer is corrected more effectively with a targeted instruction than with a new prompt. Say what is wrong and what you expect: “redo it in three bullets of no more than twelve words,” “use a more direct tone, without pleasantries,” “add the three most concrete risks.” If the result deteriorates over successive turns, the context window has often filled up: start over with a summary.
#What changes with a local model
- More explicit instructions
- A small model fills in unstated details less reliably. Write everything down, including what seems obvious to you.
- Truncated context by default
- Ollama starts at around 4,000 tokens with 24 GiB of VRAM: a long pasted text may be truncated. Increase the context window before blaming the prompt.
- Reasoning models
- They think before answering: the response arrives later and consumes reasoning tokens.
- More fragile French
- Smaller models make more agreement errors. Add a proofreading instruction or choose a larger model.
- Temperature
- A low setting (0 to 0.3) works for extraction and JSON; use a higher setting for writing. See the guide to parameters.
#A checklist before sending a prompt
- One sentence, one verb
- The task is phrased as a complete instruction, with a precise action verb.
- Separate data
- The text to be processed is enclosed in tags and never mixed with the instructions.
- Written format
- Length, structure, language, and tone are specified.
- Explicit outcome
- The model knows what to answer when information is missing.
- One or two examples
- For extraction or classification, one example is worth ten explanations.
- Verified window
- The prompt and expected response fit within the allocated context.
How Do You Query an LLM Effectively?+
Does a longer prompt produce better answers?+
Should you say “you’re an expert” in the prompt?+
How can I get reliable JSON output from a local model?+
Why does my local model ignore part of my prompt?+
What’s the difference between a system prompt and a user prompt?+
#Go further
- Master system prompts
- Temperature, top-p, top-k: the parameters
- Understanding the context window
- Function calling and JSON outputs with Ollama
- Hallucinations: how to limit them
- Your first local conversation
- Source: Anthropic, prompt engineering best practices
- Source: Ollama, structured outputs
- Source: Ollama, reasoning models
- Source: Ollama, context length
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.