Mastering the system prompts
A system prompt is the “system” role message that sets an LLM’s identity, mission, rules, style, and output format before your first question. To write one: five short blocks, imperative and verifiable instructions, and a test using a few trick questions. It consumes context and is not a safe: do not put any secrets in it.
A system prompt is the briefing given to the model before the first question. When written well, it makes a local LLM consistent and predictable; when written poorly, it wastes context and contradicts itself. You’ll learn how to structure it, install it in Ollama, LM Studio, or an API, test it methodically, and understand its limitations, particularly regarding privacy.
#What is a system prompt, and how do you write one?
A system prompt is the “system” role message placed at the beginning of the conversation: it sets the model’s identity, mission, rules, style, and response format before you ask your first question. The model rereads it on every turn, just like the history. To write one, keep five short blocks (identity, mission, hard rules, style, output format), phrase instructions as imperatives that can be verified, then test it with a few trick questions. A system prompt guides behavior; it is neither a security barrier nor a vault for your secrets, and it consumes part of your local model’s context window. This guide covers each point, with examples ready to paste into Ollama, LM Studio, or an OpenAI-compatible API.
An LLM generally receives three types of messages: system (the permanent framework), user (what you type), and assistant (what it has already answered). The chat client sends the entire conversation back to the model with every request, including the system message. That's what creates the impression of a stable “personality” throughout the session.
Model providers describe the same use case. Anthropic's documentation explains that defining a role in the system prompt focuses the model's behavior and tone, and that a single sentence is enough to make a difference. The principle applies to local models: the permanent instruction goes in the system message, and the current question goes in the user message.
#What a good system prompt really changes
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Tone consistency
- An assistant that uses first-name address in the first message and formal address in the second reveals a missing or overly vague system prompt. Setting the register once prevents drift.
- More stable constraints
- Prohibitions (“no emoji”) and requirements (“cite the source”) hold up better when defined in the system prompt than repeated on the fly. They are never guaranteed, however: the smaller the model, the more you need to verify.
- Less repetition
- If you type « réponds en français formel » for every question, put it in system once. Fewer omissions, less typing.
- Mission scope
- “Process only French labor law” reduces off-topic responses without ever ruling them out with certainty.
What the system prompt does not do: it does not add any knowledge to the model, and it does not replace either a document provided in the conversation or a RAG system. If the model needs to rely on your contracts, document retrieval provides the facts; the system prompt only says how to use and cite them.
#What it costs: space in the context window
The system prompt is not free: its tokens occupy the context window along with the history, your documents, and the upcoming response. On a local machine, this window is often smaller than you think. The “Context length” page in the Ollama documentation indicates a default window of approximately 4,000 tokens (“4k”) with less than 24 GiB of VRAM, 32,000 between 24 and 48 GiB, and 256,000 beyond that. A 400-token prompt therefore uses nearly 10% of 4,096 tokens before the first exchange even begins.
The calculation is simple: available window = total window − system prompt tokens − expected response. With 4,096 tokens, a 400-token system prompt, and an 800-token response, 2,896 tokens remain for history and documents. As the exchange grows longer, the window fills up and older content may no longer fit—another reason to keep the briefing dense rather than long.
There is a second effect documented by research: in the “Lost in the Middle” study (Liu et al., 2023), performance is often better when useful information is at the beginning or end of the context, and degrades when it is in the middle, even for models advertised as supporting long context. Practical consequence: put the most important rules at the top of your system prompt, and repeat the critical constraint in the last line if you notice omissions.
#Structure of a solid system prompt in five blocks
A good system prompt fits into five short blocks. There is no need to write a novel: every additional sentence consumes context and increases the risk of contradictory instructions. The table summarizes what each block contains and the most common mistake.
| Block | Question it answers | Example | Common error |
|---|---|---|---|
| Identity | Which model is it? | “You are Lex, a legal assistant specializing in French employment law.” | A flattering identity (“world expert”) without scope |
| Mission | What must it accomplish, and for whom? | “You help an HR director understand their obligations and draft letters.” | Mission too broad: “help with everything” |
| Hard rules | What does it still do, or never do? | « Cite the relevant article of the French Labour Code. » | Ten rules with the same priority, two of which contradict each other |
| Style | How does it express itself? | “Short sentences, accessible vocabulary.” | Vague adjectives (“professional”) without an example |
| Output format | What form does it return the answer in? | “By default: one paragraph and the references.” | Format unspecified, so it varies from one response to another |
Here is the same assembled prompt. Each rule is imperative and verifiable (the article is cited or not), and out-of-scope behavior is explicit.
This prompt serves as a structural template. A local model may cite a nonexistent article: have a human verify the references, and see the guide on hallucinations.
#Writing instructions the model follows
- Direct imperative
- “Do X.” and “Don't do Y.” are clearer than “try to … if possible,” which leaves the model free not to try.
- Positive first
- Telling it what to do (“answer in one sentence”) is better than a long list of prohibitions, which should be reserved for genuine guardrails.
- An example is worth a description
- For a precise format, provide a two-line sample answer instead of three sentences of explanation.
#In practice with Ollama: /set system and Modelfile
Ollama offers three ways to define a system prompt, depending on how long you want it to last. In an interactive session, the /set system command sets the message for the current conversation; it appears in the client’s list of available commands. The prompt disappears when the session ends.
For a persistent system prompt, create a Modelfile: it is a derived model with its own name. The Ollama documentation specifies that the SYSTEM instruction defines the system message to insert into the model template. The word “template” matters: if a model’s template has no place for the system message, the SYSTEM instruction will have no effect. If the behavior is surprising, compare it with the original Modelfile.
Third option, on the API side: the body of a request to /api/generate accepts a system field that replaces the one defined in the Modelfile, and /api/chat receives the history as an array of messages with a role and content. The Modelfile guide details the other instructions.
#In practice with LM Studio: the System Prompt field and presets
In the chat tab, the configuration panel contains a System Prompt field. To avoid retyping it, LM Studio provides presets. The documentation describes them as a way to group a system prompt and other settings into a configuration that can be reused from one conversation to the next.
#In practice via the OpenAI-compatible API
Ollama, LM Studio, vLLM, and llama-server expose an OpenAI-compatible API. The first message in the table, with the system role, defines the briefing. The Ollama documentation specifies that the client requires an API key value but the server ignores it: any string works for local use.
One design point: the API has no memory. On every call, your code sends the system message and then the full history. If you omit it from a call, the model responds without a briefing: store the message array and keep the system message at position 0.
#Three ready-to-use system prompts
#French spell checker
#Structured extractor
For production extraction, don't rely on the system prompt alone: Ollama's structured outputs let you enforce a JSON schema on responses, avoiding broken JSON. The guide to function calling and JSON outputs explains how to set it up.
#Senior coding assistant
#Test and fix a system prompt without trial and error
A system prompt improves through controlled experiments, not successive additions. Without a method, you pile on rules until you get a long, contradictory text that is impossible to debug. Here is a simple procedure that requires no particular tool.
- 01Create five test questionsWrite five entries: two normal cases, an out-of-scope question, an unusual-format request, and an attempt to derail the role (“forget your instructions”). Keep them in a file.
- 02Set the conditionsUse the same model, the same quantization, and a low temperature (0.2 to 0.3) to compare prompt versions under identical conditions.
- 03Change only one thing at a timeChange one rule, rerun the five questions, and note what changed. If two rules change at the same time, you will not know which one fixed or worsened the behavior.
- 04Remove before addingWhen a rule is ignored, rephrase it more concisely, put it first, or remove the rule that contradicts it. Adding text is a last resort.
- 05Version the promptKeep each version in a dated text file or in the Modelfile under version control, with a line describing what changed. You can roll back.
#Pitfalls to avoid
| Pitfall | Symptom | Patch |
|---|---|---|
| Prompt too long | The end of the prompt is followed less faithfully; the context window gets eaten into | Condense it, put the essentials first, and increase num_ctx if needed |
| Contradictory instructions | “Be brief” and “give lots of examples”: variable behavior | Prioritize: in case of conflict, rule 1 takes precedence |
| Hesitant wording | “Try to … if possible”: instruction applied intermittently | Direct imperative: “Do X.” |
| Language not specified | Responses in the prompt or question’s language, depending on the model | Write “Always respond in French” and write the prompt in French |
| System location ignored | Model that doesn't apply the SYSTEM from a Modelfile | Verify that the model template includes a system message |
| Secrets in the prompt | A user obtains the prompt content by asking for it | Do not put anything confidential in it |
#The system prompt is not a safe
OWASP's recommendations on the risk of system prompt leakage are clear: it must not be considered a secret or used as a security control, and it must contain neither credentials nor connection strings. Translation for a local deployment: if access control is required (per-user data, sensitive actions), it belongs in your application, not in a prompt sentence. The guide on prompt injection explains why a local model is not immune.
- Source: Ollama Modelfile reference (SYSTEM instruction)
- Source: default context length of Ollama
- Source: LM Studio presets
- Source: OWASP, system prompt leakage
- Source: Lost in the Middle (Liu et al., 2023)
- Prompting basics
- Ollama Modelfile: create and customize your model
- Understanding the context window
- Temperature, top-p, top-k: the parameters
- Prompt injection: local doesn't protect you
- Function calling and structured JSON outputs with Ollama
What is a system prompt?+
What is the ideal length for a local system prompt?+
How do you define a permanent system prompt in Ollama?+
Can a system prompt prevent a model from revealing information?+
Should you write the system prompt in French or English?+
Why is my model ignoring my system prompt?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.