Intermediate 13 minPrompting

Mastering the system prompts

Direct response

A system prompt is the “system” role message that sets an LLM’s identity, mission, rules, style, and output format before your first question. To write one: five short blocks, imperative and verifiable instructions, and a test using a few trick questions. It consumes context and is not a safe: do not put any secrets in it.

A system prompt is the briefing given to the model before the first question. When written well, it makes a local LLM consistent and predictable; when written poorly, it wastes context and contradicts itself. You’ll learn how to structure it, install it in Ollama, LM Studio, or an API, test it methodically, and understand its limitations, particularly regarding privacy.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What is a system prompt, and how do you write one?

A system prompt is the “system” role message placed at the beginning of the conversation: it sets the model’s identity, mission, rules, style, and response format before you ask your first question. The model rereads it on every turn, just like the history. To write one, keep five short blocks (identity, mission, hard rules, style, output format), phrase instructions as imperatives that can be verified, then test it with a few trick questions. A system prompt guides behavior; it is neither a security barrier nor a vault for your secrets, and it consumes part of your local model’s context window. This guide covers each point, with examples ready to paste into Ollama, LM Studio, or an OpenAI-compatible API.

An LLM generally receives three types of messages: system (the permanent framework), user (what you type), and assistant (what it has already answered). The chat client sends the entire conversation back to the model with every request, including the system message. That's what creates the impression of a stable “personality” throughout the session.

i
Analogy
A system prompt is to an LLM what a job description is to a temp worker. Without one, it does its best. With one, it knows what’s expected, what’s prohibited, and what format to use for its work.

Model providers describe the same use case. Anthropic's documentation explains that defining a role in the system prompt focuses the model's behavior and tone, and that a single sentence is enough to make a difference. The principle applies to local models: the permanent instruction goes in the system message, and the current question goes in the user message.

#What a good system prompt really changes

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Tone consistency
An assistant that uses first-name address in the first message and formal address in the second reveals a missing or overly vague system prompt. Setting the register once prevents drift.
More stable constraints
Prohibitions (“no emoji”) and requirements (“cite the source”) hold up better when defined in the system prompt than repeated on the fly. They are never guaranteed, however: the smaller the model, the more you need to verify.
Less repetition
If you type « réponds en français formel » for every question, put it in system once. Fewer omissions, less typing.
Mission scope
“Process only French labor law” reduces off-topic responses without ever ruling them out with certainty.

What the system prompt does not do: it does not add any knowledge to the model, and it does not replace either a document provided in the conversation or a RAG system. If the model needs to rely on your contracts, document retrieval provides the facts; the system prompt only says how to use and cite them.

#What it costs: space in the context window

The system prompt is not free: its tokens occupy the context window along with the history, your documents, and the upcoming response. On a local machine, this window is often smaller than you think. The “Context length” page in the Ollama documentation indicates a default window of approximately 4,000 tokens (“4k”) with less than 24 GiB of VRAM, 32,000 between 24 and 48 GiB, and 256,000 beyond that. A 400-token prompt therefore uses nearly 10% of 4,096 tokens before the first exchange even begins.

The calculation is simple: available window = total window − system prompt tokens − expected response. With 4,096 tokens, a 400-token system prompt, and an 800-token response, 2,896 tokens remain for history and documents. As the exchange grows longer, the window fills up and older content may no longer fit—another reason to keep the briefing dense rather than long.

There is a second effect documented by research: in the “Lost in the Middle” study (Liu et al., 2023), performance is often better when useful information is at the beginning or end of the context, and degrades when it is in the middle, even for models advertised as supporting long context. Practical consequence: put the most important rules at the top of your system prompt, and repeat the critical constraint in the last line if you notice omissions.

→
Increase the window if necessary
If your system prompt exceeds a few hundred tokens, or if you attach documents, increase the context window rather than shortening it blindly: PARAMETER num_ctx in a Modelfile, or the OLLAMA_CONTEXT_LENGTH variable when starting the server. A larger window uses more video memory; the guide to the context window explains the trade-off.

#Structure of a solid system prompt in five blocks

A good system prompt fits into five short blocks. There is no need to write a novel: every additional sentence consumes context and increases the risk of contradictory instructions. The table summarizes what each block contains and the most common mistake.

The five blocks of a system prompt
BlockQuestion it answersExampleCommon error
IdentityWhich model is it?“You are Lex, a legal assistant specializing in French employment law.”A flattering identity (“world expert”) without scope
MissionWhat must it accomplish, and for whom?“You help an HR director understand their obligations and draft letters.”Mission too broad: “help with everything”
Hard rulesWhat does it still do, or never do?« Cite the relevant article of the French Labour Code. »Ten rules with the same priority, two of which contradict each other
StyleHow does it express itself?“Short sentences, accessible vocabulary.”Vague adjectives (“professional”) without an example
Output formatWhat form does it return the answer in?“By default: one paragraph and the references.”Format unspecified, so it varies from one response to another

Here is the same assembled prompt. Each rule is imperative and verifiable (the article is cited or not), and out-of-scope behavior is explicit.

Full system prompt
Tu es Lex, assistant juridique spécialisé en droit
du travail français (version 2026).

MISSION
Tu aides un DRH à comprendre ses obligations et à
rédiger des courriers RH conformes.

RÈGLES
- Ne jamais donner d'avis juridique formel.
- Toujours citer l'article du Code du travail concerné,
  par ex. "L.1232-2".
- Si la question sort du droit du travail, répondre :
  "Hors de mon périmètre. Je ne traite que le droit
  du travail français."

STYLE
- Phrases courtes, vocabulaire accessible.
- Pas d'anglicisme si un équivalent existe.
- Ton professionnel mais direct.

FORMAT
- Réponse par défaut : 1 paragraphe + références.
- Sur demande : tableau, courrier-type, checklist.

This prompt serves as a structural template. A local model may cite a nonexistent article: have a human verify the references, and see the guide on hallucinations.

#Writing instructions the model follows

Direct imperative
“Do X.” and “Don't do Y.” are clearer than “try to … if possible,” which leaves the model free not to try.
Positive first
Telling it what to do (“answer in one sentence”) is better than a long list of prohibitions, which should be reserved for genuine guardrails.
An example is worth a description
For a precise format, provide a two-line sample answer instead of three sentences of explanation.

#In practice with Ollama: /set system and Modelfile

Ollama offers three ways to define a system prompt, depending on how long you want it to last. In an interactive session, the /set system command sets the message for the current conversation; it appears in the client’s list of available commands. The prompt disappears when the session ends.

Interactive session
>>> /set system """
Tu es Lex, assistant juridique spécialisé en droit
du travail français...
"""

For a persistent system prompt, create a Modelfile: it is a derived model with its own name. The Ollama documentation specifies that the SYSTEM instruction defines the system message to insert into the model template. The word “template” matters: if a model’s template has no place for the system message, the SYSTEM instruction will have no effect. If the behavior is surprising, compare it with the original Modelfile.

Modelfile
FROM mistral

PARAMETER temperature 0.3
PARAMETER num_ctx 8192

SYSTEM """
Tu es Lex, assistant juridique...
"""
Terminal
ollama create lex -f ./Modelfile
ollama run lex

Third option, on the API side: the body of a request to /api/generate accepts a system field that replaces the one defined in the Modelfile, and /api/chat receives the history as an array of messages with a role and content. The Modelfile guide details the other instructions.

#In practice with LM Studio: the System Prompt field and presets

In the chat tab, the configuration panel contains a System Prompt field. To avoid retyping it, LM Studio provides presets. The documentation describes them as a way to group a system prompt and other settings into a configuration that can be reused from one conversation to the next.

→
A preset library
Every parameter in the Advanced Configuration panel can be saved in a named preset: system prompt, as well as temperature, top-p, or maximum number of tokens. Create one preset per use case (“FR proofreader,” “code assistant,” “legal expert”) and change roles with one click. Presets can also be imported from a file or URL.

#In practice via the OpenAI-compatible API

Ollama, LM Studio, vLLM, and llama-server expose an OpenAI-compatible API. The first message in the table, with the system role, defines the briefing. The Ollama documentation specifies that the client requires an API key value but the server ignores it: any string works for local use.

Python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # ignoré
)

resp = client.chat.completions.create(
    model="mistral",
    messages=[
        {"role": "system", "content": "Tu es Lex..."},
        {"role": "user",   "content": "Peut-on rompre une période d'essai ?"},
    ],
)

print(resp.choices[0].message.content)

One design point: the API has no memory. On every call, your code sends the system message and then the full history. If you omit it from a call, the model responds without a briefing: store the message array and keep the system message at position 0.

#Three ready-to-use system prompts

#French spell checker

System prompt
Tu es un correcteur de français académique.
Règle unique : renvoie UNIQUEMENT le texte corrigé,
sans commentaire, sans liste des changements.
Corrige : orthographe, grammaire, ponctuation, barbarismes.
Ne reformule jamais. Ne change pas le style de l'auteur.

#Structured extractor

System prompt
Tu extrais des informations structurées depuis du texte libre.
Sortie STRICTE : JSON valide, rien d'autre (pas de ```json, pas de texte).
Si une information est absente, renvoie null.
Ne devine jamais.

For production extraction, don't rely on the system prompt alone: Ollama's structured outputs let you enforce a JSON schema on responses, avoiding broken JSON. The guide to function calling and JSON outputs explains how to set it up.

#Senior coding assistant

System prompt
Tu es un ingénieur logiciel senior.
Pour toute réponse de code :
1. Dis en une phrase ce que fait le code.
2. Donne le code dans un bloc unique.
3. Liste 1 à 3 cas limites à surveiller.
Tu ne mets JAMAIS de commentaire à l'intérieur du code,
sauf si un piège non-évident le justifie.

#Test and fix a system prompt without trial and error

A system prompt improves through controlled experiments, not successive additions. Without a method, you pile on rules until you get a long, contradictory text that is impossible to debug. Here is a simple procedure that requires no particular tool.

  1. 01
    Create five test questions
    Write five entries: two normal cases, an out-of-scope question, an unusual-format request, and an attempt to derail the role (“forget your instructions”). Keep them in a file.
  2. 02
    Set the conditions
    Use the same model, the same quantization, and a low temperature (0.2 to 0.3) to compare prompt versions under identical conditions.
  3. 03
    Change only one thing at a time
    Change one rule, rerun the five questions, and note what changed. If two rules change at the same time, you will not know which one fixed or worsened the behavior.
  4. 04
    Remove before adding
    When a rule is ignored, rephrase it more concisely, put it first, or remove the rule that contradicts it. Adding text is a last resort.
  5. 05
    Version the prompt
    Keep each version in a dated text file or in the Modelfile under version control, with a line describing what changed. You can roll back.

#Pitfalls to avoid

Common errors and fixes
PitfallSymptomPatch
Prompt too longThe end of the prompt is followed less faithfully; the context window gets eaten intoCondense it, put the essentials first, and increase num_ctx if needed
Contradictory instructions“Be brief” and “give lots of examples”: variable behaviorPrioritize: in case of conflict, rule 1 takes precedence
Hesitant wording“Try to … if possible”: instruction applied intermittentlyDirect imperative: “Do X.”
Language not specifiedResponses in the prompt or question’s language, depending on the modelWrite “Always respond in French” and write the prompt in French
System location ignoredModel that doesn't apply the SYSTEM from a ModelfileVerify that the model template includes a system message
Secrets in the promptA user obtains the prompt content by asking for itDo not put anything confidential in it

#The system prompt is not a safe

OWASP's recommendations on the risk of system prompt leakage are clear: it must not be considered a secret or used as a security control, and it must contain neither credentials nor connection strings. Translation for a local deployment: if access control is required (per-user data, sensitive actions), it belongs in your application, not in a prompt sentence. The guide on prompt injection explains why a local model is not immune.

FAQ
What is a system prompt?+
This is the « system » role message placed at the beginning of a conversation with an LLM. It defines the model's identity, mission, rules, style, and response format. It remains in the context on every turn, stabilizing behavior throughout the session without requiring you to repeat your instructions.
What is the ideal length for a local system prompt?+
Keep it as short as possible while covering your rules: a few lines for a simple case, one to two hundred words for a business assistant. Every token used by the system reduces the space available for history, especially with the default Ollama window of about 4,000 tokens under 24 GiB of VRAM. Put the essentials first.
How do you define a permanent system prompt in Ollama?+
Create a Modelfile with a FROM line and a SYSTEM instruction containing your text, then run ollama create with the desired name and ollama run with that name. The prompt is then embedded in the derived model. The /set system command applies only to the current interactive session.
Can a system prompt prevent a model from revealing information?+
Not reliably. OWASP reminds us that the system prompt should not be considered a secret or used as a security control. A motivated user can extract it or subvert the instruction. Put access controls and sensitive data in your application, never in the prompt text.
Should you write the system prompt in French or English?+
Write it in the language expected for the responses, and add an explicit instruction such as « Always respond in French ». Some models respond in the prompt’s language rather than the question’s language. Test with two or three questions in each language on your specific model before generalizing.
Why is my model ignoring my system prompt?+
First, verify that the prompt is being sent correctly (the system role first, or SYSTEM in the Modelfile) and that the model's template supports a system message. Then shorten it, remove contradictory instructions, switch to the imperative, and test with a low temperature. A very small model follows complex instructions less reliably.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.