Intermediate 12 minOllama

Ollama Modelfile: create and customize your model

Ollama is not limited to downloading existing models: with a text file a few lines long, you can build your own variants. A Modelfile is the LLM equivalent of a Dockerfile—it freezes a system prompt, sampling parameters, and a conversation template. This guide shows how to customize Ollama with a Modelfile through three concrete use cases: a strict French assistant, a Python coder, and a no-chatter translator.

By Mohamed Meguedmi·Update 2026-08-31·Tested on Windows, macOS, and Linux
i
In brief
An Ollama Modelfile is the equivalent of a Dockerfile for an LLM: a text file that locks in behavior on top of an existing model without touching its weights. · Three directives cover the essentials: FROM (the base model), SYSTEM (the persistent system prompt), and PARAMETER (temperature, context, etc.). · Once the file is written, ollama create mon-modele -f ./Modelfile creates the variant, which can then be used as ollama run mon-modele. · Useful for a strictly French-speaking assistant, a silent Python coder, or a no-chatter translator.

#Why use a Modelfile?

When you type ollama run qwen3.5:9b, you get a generic model. It sometimes answers a question in French in English, comments on its code answers with endless introductions, and negotiates when you ask for a literal translation. Instead of repeating the same system prompt in every session, you freeze it once in a Modelfile and get a reusable variant.

In practice, a Modelfile does three things: (1) attaches a persistent system prompt to the model, (2) configures generation parameters (temperature, context, stop tokens), and (3) optionally replaces the chat template. The created variant is used like any Ollama model: ollama run mon-assistant.

i
This isn’t fine-tuning
A Modelfile does not modify the model's weights. It configures the model's behavior on top of them. To truly teach a model a style or domain, you need fine-tuning (LoRA, QLoRA) — which goes beyond Ollama alone.

#Prerequisites

The Local AI Kit

You know how to customize a model with a Modelfile. The Local AI Kit installs it in a real private ChatGPT for the whole household, with Ollama and Open WebUI (ch. 6), and helps you choose the base model suited to your machine (ch. 3).

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Ollama installed
Version 0.3+ recommended. Check with ollama --version. The daemon must be running on http://localhost:11434.
An already pulled base model
For example, ollama pull qwen3.5:9b or ollama pull qwen3-coder:30b. The Modelfile inherits from this model.
A text editor
VSCode, Notepad++, vim — any of them. The file does not require an extension; by convention, it is generally called Modelfile.
Approximately 5 minutes
It takes just long enough to create the file and launch ollama create. No additional download: an already-present model is reused.

#1. Modelfile syntax

A Modelfile is a text file with uppercase instructions, one per block. The order does not matter, but convention places FROM first.

Minimal skeleton
FROM qwen3.5:9b

SYSTEM """
Tu es un assistant utile et concis.
"""

PARAMETER temperature 0.7
PARAMETER num_ctx 4096

The main instructions:

FROM
Base model. Required. Accepts a Ollama name (qwen3.5:9b, qwen3.8:27b) or a local path to a GGUF.
SYSTEM
The system prompt injected into every conversation. Use triple quotes for multiline text.
PARAMETER
Sampling and context settings. One instruction per parameter. The most useful ones: temperature, top_p, top_k, num_ctx, repeat_penalty, num_predict, stop.
TEMPLATE
Full prompt format (rare). Only change it if you know what you're doing—each model family has its own template.
MESSAGE
Few-shot examples. Prefixed with MESSAGE user or MESSAGE assistant. Useful for anchoring a style.
ADAPTER
Path to a GGUF LoRA to apply on top of FROM. For advanced users.
LICENSE
Free text, for traceability. Does not affect execution.
→
The parameters that really matter
For 90% of use cases, you only need SYSTEM + temperature + num_ctx. No need to configure everything: Ollama applies sensible defaults (temperature 0.8, num_ctx 2048, repeat_penalty 1.1).

#2. Create your first model

The workflow is always the same: write a Modelfile, run ollama create, and test with ollama run.

  1. 01
    Create the file
    In a working directory, create a file named Modelfile (with no extension). Put your variant definition in it.
  2. 02
    Start creation
    From the same folder: ollama create mon-modele -f ./Modelfile. Ollama validates the syntax, inherits the model weights from FROM, and registers a new entry. It is instantaneous and does not download anything again.
  3. 03
    Check
    ollama list montre votre nouveau modèle aux côtés des autres. La taille affichée est identique à celle du modèle de base — c'est une référence, pas une copie.
  4. 04
    Test
    ollama run mon-modele. Le system prompt et les paramètres sont déjà appliqués. Vous pouvez aussi pointer une interface (Open WebUI, LM Studio) sur cette variante via l'API.
Complete cycle
# Dans le dossier qui contient le Modelfile
ollama create assistant-fr -f ./Modelfile
ollama list
ollama run assistant-fr
!
The model name is final
To modify a Modelfile, edit the file and then rerun ollama create with the SAME name — the variant is replaced. If you change the name, you accumulate references. ollama rm <name> cleans up a variant that is no longer needed.

#3. Example: concise French assistant

Goal: an assistant that consistently responds in French, without preambles such as "Of course, here’s…", without apologetic formulas, and with calibrated length.

Modelfile-assistant-fr
FROM qwen3.5:9b

SYSTEM """
Tu es un assistant francophone précis et concis.

Règles strictes :
- Réponds toujours en français, jamais en anglais, même si l'utilisateur écrit en anglais.
- Pas de préambule ("Bien sûr", "Voici", "Absolument"). Va directement au fait.
- Pas d'excuses ni de disclaimers sauf si l'information est réellement incertaine.
- Réponses courtes par défaut (3-5 phrases). Détaille uniquement si on te le demande explicitement.
- Si tu ne sais pas, dis-le en une phrase. N'invente pas.
"""

PARAMETER temperature 0.6
PARAMETER top_p 0.9
PARAMETER num_ctx 8192
PARAMETER repeat_penalty 1.15

Creation and testing:

Terminal
ollama create assistant-fr -f ./Modelfile-assistant-fr
ollama run assistant-fr "Quels sont les avantages d'un LLM local ?"

A repeat_penalty of 1.15 (instead of the default 1.1) stops models from repeating themselves. A num_ctx of 8192 lets you paste somewhat longer documents without breaking the default 2048-token context.

→
Test the effect of the system prompt
Compare ollama run qwen3.5:9b and ollama run assistant-fr side by side on the same question. Without a Modelfile, Qwen 3.5 often starts with “Sure,”. With your variant, it goes straight to the answer. That’s the practical difference.

#4. Example: quiet Python coding

Goal: a coding companion that returns Python ready to paste, with minimal surrounding prose. Ideal for workflows where you want to pipe the output into an editor or script.

Modelfile-coder-py
FROM qwen3-coder:30b

SYSTEM """
Tu es un assistant de code Python expert. Tu suis ces règles :

1. Réponds uniquement avec du code Python valide, dans un bloc fenced ```python.
2. Pas d'explication avant ou après le code, sauf si l'utilisateur demande explicitement une explication.
3. Le code doit être complet, exécutable, avec les imports nécessaires en haut.
4. Privilégie la stdlib quand c'est possible. Si une dépendance externe est nécessaire, mets un commentaire en tête : # pip install <package>.
5. Type hints quand c'est raisonnable (Python 3.10+).
6. Si la demande est ambiguë, choisis l'interprétation la plus probable et code-la, sans poser de questions.
"""

PARAMETER temperature 0.2
PARAMETER top_p 0.95
PARAMETER num_ctx 16384
PARAMETER stop "```\n\n"

A few notable technical choices:

FROM qwen3-coder:30b
A code-specialized model from 2026 (MoE 30B-A3B, only 3B active parameters, 256k context, ~19 GB in Q4). On a 24 GB card (RTX 4090) or a 32 GB Mac, it runs at 50–80 tokens/s thanks to its 3B active parameters. If you only have 16 GB of VRAM, replace it with gpt-oss:20b (14 GB) or devstral:24b.
temperature 0.2
Low for coding. At 0.7 (default), the model improvises variable names and sometimes invents APIs. For code, you want determinism.
num_ctx 16384
Extended window for analyzing entire files. Warning: it multiplies the VRAM consumed by the context.
stop
The model stops after the first code block. This avoids post-code explanations cluttering the output.
Usage example
ollama create coder-py -f ./Modelfile-coder-py
ollama run coder-py "Lire un CSV avec pandas et retourner la moyenne de la colonne 'prix'"
i
Adjust for your needs
If you also want pytest tests generated systematically, add a rule to SYSTEM: "7. Also produce a minimal pytest test block after the main code." The model will follow the instruction as long as it is clear and non-contradictory.

#5. Example: strict EN→FR translator

Goal: a translator that returns ONLY the translation, without quotation marks, without "Translation:", and without stylistic variation. The typical use case: integration into an n8n pipeline or a bash script where the output will be consumed raw.

Modelfile-translator
FROM qwen3.5:9b

SYSTEM """
Tu es un traducteur professionnel anglais → français.

Règles ABSOLUES :
- Tu reçois un texte en anglais. Tu retournes UNIQUEMENT la traduction française.
- Pas de préfixe ("Traduction :", "Voici :"), pas de guillemets autour de ta réponse, pas de commentaire.
- Conserve la mise en forme exacte du texte source (sauts de ligne, listes, ponctuation).
- Conserve les noms propres, marques, URLs et identifiants techniques tels quels.
- Si le texte est déjà en français, renvoie-le inchangé.
- Si le texte est dans une autre langue que l'anglais, renvoie : ERREUR_LANGUE_NON_SUPPORTÉE
"""

PARAMETER temperature 0.3
PARAMETER num_ctx 4096
PARAMETER num_predict 2048

Adding MESSAGE to anchor the format through few-shot prompting makes the behavior even more predictable:

Enhanced version with few-shot
FROM qwen3.5:9b

SYSTEM """
Tu traduis de l'anglais vers le français. Tu renvoies UNIQUEMENT la traduction, sans préfixe ni guillemets.
"""

MESSAGE user "The meeting starts at 3 PM."
MESSAGE assistant "La réunion commence à 15 heures."

MESSAGE user "Please check the attached document."
MESSAGE assistant "Merci de consulter le document ci-joint."

PARAMETER temperature 0.3

Pipeline usage:

Shell pipe
echo "The deployment is scheduled for tomorrow morning." | ollama run translator-en-fr
→
Direct call through the API
To integrate it into a script, the HTTP API is more convenient: curl http://localhost:11434/api/generate -d '{"model":"translator-en-fr","prompt":"Your text here","stream":false}'. The JSON response contains the translation in the response field.

#6. Managing custom models

Over time, you'll accumulate several variants. A few useful commands to stay in control:

Inventory and cleanup
# Lister tous les modèles (originaux + variantes)
ollama list

# Voir le Modelfile d'une variante (utile pour récupérer un Modelfile perdu)
ollama show --modelfile assistant-fr

# Voir les paramètres effectifs
ollama show --parameters assistant-fr

# Voir le system prompt seul
ollama show --system assistant-fr

# Supprimer une variante (libère uniquement l'entrée, pas les poids du modèle de base)
ollama rm assistant-fr
i
Version your Modelfiles in Git
Modelfiles are text. Put them in a separate Git repo. You keep the history of prompts that work, and any machine with Ollama and the right base model can recreate all your variants in a few seconds.

#7. Which base model should you choose?

Every Modelfile starts with FROM: choosing the base model determines 90% of the final result. A perfect system prompt can never make up for a base that is poorly suited to your use case or too large for your VRAM.

General-purpose assistant
A recent 8B to 14B model (from the Qwen or Gemma family) in Q4: responsive, good in French, and leaves room for context.
Code
A code-specialized model—MoE models such as Qwen3-Coder 30B-A3B are very fast with 20 GB of memory or more; below that, a 7–9B code model remains a good choice.
Formal French / writing
Models Mistral (Nemo, Magistral) retain the advantage in French at the same size.
Small configuration (8 GB of VRAM)
Stick to 4B to 9B in Q4: a model that spills into RAM ruins responsiveness, regardless of the Modelfile.

To compare models by VRAM, license, and use case, the QuelLLM catalog filters models compatible with your machine—each entry provides the ollama run exacte command to put in your FROM.

#Troubleshooting

The model ignores the system prompt
Make sure you're actually launching your variant (ollama run mon-modele) and not the base model. Confirm with ollama show --system mon-modele.
Truncated responses
Increase num_predict (default 128 in some configs) to 2048 or higher. Or add PARAMETER num_predict -1 to disable the limit.
The model does not follow a strict rule
Restate the rule in uppercase, add "ABSOLUTELY" or "NEVER", and add an example via MESSAGE. Small models (under ~10B) struggle with isolated negations.
Error "Error: invalid model reference"
The FROM model is not installed locally. Run ollama pull <modele> before the ollama create.
VRAM saturated after creation
A high num_ctx multiplies consumption. If you go from 2048 to 16384, expect +1 to +3 GB of VRAM depending on the model size. Drop back to 8192 if memory gets tight.

#Go further

The Modelfile is the basic building block for specializing Ollama without touching the weights. Three natural directions to explore next:

Mastering system prompts in greater depth
The guide to system prompts explains how to structure roles, constraints, and examples for truly reliable behavior—useful as soon as you exceed 5 lines of SYSTEM.
Adjust temperature, top-p, and top-k
Before touching the model, knowing precisely what each sampling parameter changes avoids a lot of back-and-forth. The dedicated guide covers all useful settings.
Expose your variants through Open WebUI
Once you create your custom models, Open WebUI lists them automatically. You can switch between assistant-fr, coder-py, and translator-en-fr with one click from the interface.
Frequently asked questions about Ollama models
How do you list the installed Ollama models?+
ollama list affiche tous les modèles présents sur le disque (nom, tag, taille, date). ollama ps montre ceux actuellement chargés en mémoire, avec la répartition CPU/GPU — c'est le premier réflexe quand une réponse semble anormalement lente.
Where are the Ollama models stored?+
In ~/.ollama/models on macOS and Linux, and in C:\Users\<you>\.ollama\models on Windows. To move them (for example, to another drive), set the OLLAMA_MODELS environment variable before starting the daemon.
How do you delete a model to free up space?+
ollama rm nom:tag supprime la référence. Les poids eux-mêmes (blobs) ne sont réellement effacés que si aucun autre modèle ne les partage — supprimer un modèle personnalisé bâti sur une base conservée ne libère donc que quelques kilo-octets.
Does a custom model take up twice as much disk space?+
No. Ollama stores weights in content-addressed blobs: your custom model references the same blobs as its base model. Only the Modelfile, system prompt, and parameters are added—create as many variants as you want; the disk cost is negligible.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.