Beginner 11 minPrompting

The basics of prompting

Direct response

To query an LLM, write a complete instruction as if you were briefing a colleague who knows neither your project nor your habits: the precise task, the relevant context, the data to process clearly separated, and the expected response format. You can send it through a chat interface, the command line, or the API. The model cannot infer your intent or the format: anything you do not write will be improvised.

This guide first answers the question in the title, then details the structure of a prompt, the errors that cost the most time, the techniques that work on almost every model, and what changes with a small local model. The principles are based on Anthropic's official prompt engineering documentation and that of Ollama.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#How to query an LLM: three ways, one rule

A language model continues text. What it produces therefore depends almost entirely on what you give it. Anthropic's documentation states a useful golden rule for any model: show your prompt to a colleague who does not know the task and ask them to execute it; if they are lost, the model will be too. It recommends explicitly asking for what you want, including the level of effort, rather than expecting the model to infer it from a vague instruction. This rule applies just as much to an 8-billion-parameter model on your computer as to a cutting-edge model. Parameters, quantization, and the chosen model matter, but a vague prompt ruins any model.

#Three ways to send a question to a local model

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Chat interface
Open WebUI, LM Studio, Jan: you type, the model responds, and the history is retained. The simplest way to get started.
Command line
ollama run suivi du nom du modèle ouvre une session interactive dans le terminal. Pratique pour des essais rapides.
API
A program sends an HTTP request to the local server and receives the response as JSON. This is the path for scripts and automations.
Query a model via the Ollama API
curl http://localhost:11434/api/chat -d '{
  "model": "gpt-oss",
  "messages": [
    {"role": "system", "content": "Tu réponds en français, en trois phrases maximum."},
    {"role": "user", "content": "Explique ce qu'est un token."}
  ],
  "stream": false
}'

The system role message sets the permanent framework, while the user role message contains the request. The guide to system prompts explains what to put in them. Whichever route you choose, the quality of the response depends on what the message contains.

#Anatomy of a good prompt: role, task, context, format

Many guides focus on four elements: role, task, context, and format. This isn't a standard, but a mnemonic that aligns with the documentation's recommendations: be precise, provide context, structure the content, set the format, and, when needed, specify a role. Anthropic's documentation notes that a one-sentence role in the system message already steers tone and behavior.

Role
A sentence that situates the model: “You are a rigorous French proofreader.” It sets the register, not the intelligence.
Task
The action verb, as precise as possible: “Summarize in five factual bullet points, without adjectives” is better than “Summarize.”
Context
The data to process and the reason for the request. Explaining why a constraint exists helps the model generalize it correctly.
Format
Length, structure, language, tone. Without these details, the model chooses for itself.
A complete prompt
Tu es un assistant juridique qui explique en français simple.

Tâche : résume le contrat ci-dessous en 5 points, du plus risqué
au moins risqué pour le signataire.

<contrat>
[le texte du contrat ici]
</contrat>

Format : liste à puces, une phrase par puce, 20 mots maximum.

The tags enclosing the contract in the example above are not decorative: the documentation states that XML tags help the model unambiguously distinguish instructions, context, examples, and variable inputs. Choose consistent, descriptive names, and use the same delimiter from one prompt to the next.

#The errors that waste the most time

Error, symptom, and fix
ErrorSymptomCorrection
Keyword prompt (“contract summary risks”)The model does not distinguish between a question and a commandWrite a complete sentence with a verb
Data mixed with instructionsThe model treats part of your instructions as contentWrap the data in tags
Implicit expectationsUnexpected length, language, or toneSpecify the length, language, and tone
Prohibitions only (“don't use Markdown”)The model does it anywayState what you want: “respond in written paragraphs”
Long prompt without structureIgnored instructionsNumber the steps, separate the blocks
Constraint without a reasonAwkward applicationExplain why: “the text will be read aloud”

On the penultimate line, Anthropic’s documentation recommends telling the model what to do rather than what not to do, and gives the example of replacing “don’t use Markdown” with a description of the desired format: paragraphs of flowing prose. On the last line, it recommends giving the reason for an instruction: the model derives a general rule from it rather than a literal prohibition.

#Techniques that work on almost all models

#Give examples

According to Anthropic’s documentation, examples are one of the most reliable ways to guide format, tone, and structure. Three to five examples generally produce the best results, provided they are relevant, varied (including edge cases), and enclosed in tags so they are not mistaken for instructions. This is particularly effective for classification, extraction, or reformatting.

#Ask it to reason

For a logic problem, calculation, or code, asking the model to lay out the steps before concluding often improves the result, at the cost of a longer response. Some local models do this on their own: Ollama’s documentation explains that reasoning models return a thinking field separate from the final answer, which you can read, display, or hide. For other models, one sentence is enough: “Explain your reasoning step by step, then give the final answer.”

#Place a long document

Anthropic's documentation recommends that, for inputs longer than 20,000 tokens, you place long documents at the top of the prompt, before the question, and write the question at the end: according to its tests, this can improve response quality by up to 30%, especially for inputs containing multiple documents. This result comes from tests on Claude models: verify it on your local model with your own texts.

#Have the response checked

Ask the model to cite, for every claim, the passage from the provided text that supports it, or to write “unverified” if it cannot find one. This instruction does not eliminate hallucinations, but it makes errors easier to spot. The guide on hallucinations explains its limitations in detail.

#A concrete example: from a vague instruction to a usable prompt

Take a typical request: you paste in meeting minutes and write “summarize this.” The model does not know who the summary is for, how many lines you expect, whether it should list decisions or tasks, or which language to use. It produces a generic text of arbitrary length, which you will correct manually.

The same prompt, before and after
ItemBlurry versionSpecified version
TaskGive me a summarySummarize this report in five bullet points
PublicNot specifiedFor a director who was not present
Expected contentNot specifiedDecisions made, followed by actions with an owner and date
FormatFreeBulleted list, one sentence per bullet, 20 words maximum
Garde-fouNoneIf a date or owner is missing, write “not specified” instead of making it up

The last line is the most useful with a local model: giving the model an explicit fallback when information is missing reduces hallucinations. Each addition takes only one sentence, and together they turn a result that needs reworking into a usable one. Test it again on three different reports before locking in your prompt: a prompt that succeeds on only one example is fragile. Then keep your best version in a text file: reused as-is, it will save you more time than any wording trick, and it will let you compare two models with the same instructions.

#Enforce an output format

When the response must be read by a program, a simple phrase such as “respond in JSON” is not always enough. Ollama offers structured outputs: you pass a JSON schema in the request’s format field, and the response is constrained to follow it. The documentation recommends also providing the schema in the prompt to anchor the response, lowering the temperature (for example, to 0) for more deterministic results, and defining the schema with Pydantic or Zod so it can be reused for validation. Structured outputs also work through the OpenAI-compatible API with response_format. They are not available for Ollama’s cloud models.

Schema-constrained JSON output
curl http://localhost:11434/api/chat -d '{
  "model": "gpt-oss",
  "messages": [{"role": "user", "content": "Résume ce contrat : ..."}],
  "stream": false,
  "format": {"type": "object", "properties": {"resume": {"type": "string"}, "risques": {"type": "array", "items": {"type": "string"}}}, "required": ["resume", "risques"]}
}'

#Iterate without starting from scratch

An imperfect answer is corrected more effectively with a targeted instruction than with a new prompt. Say what is wrong and what you expect: “redo it in three bullets of no more than twelve words,” “use a more direct tone, without pleasantries,” “add the three most concrete risks.” If the result deteriorates over successive turns, the context window has often filled up: start over with a summary.

#What changes with a local model

More explicit instructions
A small model fills in unstated details less reliably. Write everything down, including what seems obvious to you.
Truncated context by default
Ollama starts at around 4,000 tokens with 24 GiB of VRAM: a long pasted text may be truncated. Increase the context window before blaming the prompt.
Reasoning models
They think before answering: the response arrives later and consumes reasoning tokens.
More fragile French
Smaller models make more agreement errors. Add a proofreading instruction or choose a larger model.
Temperature
A low setting (0 to 0.3) works for extraction and JSON; use a higher setting for writing. See the guide to parameters.

#A checklist before sending a prompt

One sentence, one verb
The task is phrased as a complete instruction, with a precise action verb.
Separate data
The text to be processed is enclosed in tags and never mixed with the instructions.
Written format
Length, structure, language, and tone are specified.
Explicit outcome
The model knows what to answer when information is missing.
One or two examples
For extraction or classification, one example is worth ten explanations.
Verified window
The prompt and expected response fit within the allocated context.
FAQ
How Do You Query an LLM Effectively?+
Write a complete prompt: the precise task, relevant context, data enclosed in tags, and the expected response format. Test it by giving it to a colleague who doesn't know the subject: if they hesitate, the model will too. Add two or three examples to establish the response's style and structure.
Does a longer prompt produce better answers?+
Not automatically. What helps is precision and structure: useful context, examples, and an explicit format. A long but confusing prompt, or one that mixes instructions and data, degrades the response. Add tags and numbered steps before adding words, and test each version on the same cases.
Should you say “you’re an expert” in the prompt?+
A one-sentence role guides tone and vocabulary, according to Anthropic's documentation, but it does not add knowledge. It is more effective to specify the audience, task, and format. Prefer “explain to a beginner in three steps” over “you are an expert teacher” alone.
How can I get reliable JSON output from a local model?+
With Ollama, pass a JSON schema in the request's format field: the output is constrained to follow it. Add the schema to the prompt, set the temperature to 0, and validate the response in code. This feature is not available for Ollama cloud models, according to its documentation.
Why does my local model ignore part of my prompt?+
Three common causes: the text exceeds the context window (Ollama starts at around 4,000 tokens with 24 GiB of VRAM), the model is too small to follow several instructions at once, or instructions and data are mixed together. Check the window, break up the task, and enclose the data in tags.
What’s the difference between a system prompt and a user prompt?+
The system prompt sets a permanent framework: role, tone, rules, and format. The user prompt contains the current request. Separating the two avoids repeating instructions in every message and makes the model's behavior more stable. Ollama distinguishes them through the system and user roles of the API messages.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.