Ollama Modelfile: create and customize your model
Ollama is not limited to downloading existing models: with a text file a few lines long, you can build your own variants. A Modelfile is the LLM equivalent of a Dockerfile—it freezes a system prompt, sampling parameters, and a conversation template. This guide shows how to customize Ollama with a Modelfile through three concrete use cases: a strict French assistant, a Python coder, and a no-chatter translator.
#Why use a Modelfile?
When you type ollama run qwen3.5:9b, you get a generic model. It sometimes answers a question in French in English, comments on its code answers with endless introductions, and negotiates when you ask for a literal translation. Instead of repeating the same system prompt in every session, you freeze it once in a Modelfile and get a reusable variant.
In practice, a Modelfile does three things: (1) attaches a persistent system prompt to the model, (2) configures generation parameters (temperature, context, stop tokens), and (3) optionally replaces the chat template. The created variant is used like any Ollama model: ollama run mon-assistant.
#Prerequisites
You know how to customize a model with a Modelfile. The Local AI Kit installs it in a real private ChatGPT for the whole household, with Ollama and Open WebUI (ch. 6), and helps you choose the base model suited to your machine (ch. 3).
- Lifetime online access
- PDF + files
- Lifetime updates
- Ollama installed
- Version 0.3+ recommended. Check with ollama --version. The daemon must be running on http://localhost:11434.
- An already pulled base model
- For example, ollama pull qwen3.5:9b or ollama pull qwen3-coder:30b. The Modelfile inherits from this model.
- A text editor
- VSCode, Notepad++, vim — any of them. The file does not require an extension; by convention, it is generally called Modelfile.
- Approximately 5 minutes
- It takes just long enough to create the file and launch ollama create. No additional download: an already-present model is reused.
#1. Modelfile syntax
A Modelfile is a text file with uppercase instructions, one per block. The order does not matter, but convention places FROM first.
The main instructions:
- FROM
- Base model. Required. Accepts a Ollama name (qwen3.5:9b, qwen3.8:27b) or a local path to a GGUF.
- SYSTEM
- The system prompt injected into every conversation. Use triple quotes for multiline text.
- PARAMETER
- Sampling and context settings. One instruction per parameter. The most useful ones: temperature, top_p, top_k, num_ctx, repeat_penalty, num_predict, stop.
- TEMPLATE
- Full prompt format (rare). Only change it if you know what you're doing—each model family has its own template.
- MESSAGE
- Few-shot examples. Prefixed with MESSAGE user or MESSAGE assistant. Useful for anchoring a style.
- ADAPTER
- Path to a GGUF LoRA to apply on top of FROM. For advanced users.
- LICENSE
- Free text, for traceability. Does not affect execution.
#2. Create your first model
The workflow is always the same: write a Modelfile, run ollama create, and test with ollama run.
- 01Create the fileIn a working directory, create a file named Modelfile (with no extension). Put your variant definition in it.
- 02Start creationFrom the same folder: ollama create mon-modele -f ./Modelfile. Ollama validates the syntax, inherits the model weights from FROM, and registers a new entry. It is instantaneous and does not download anything again.
- 03Checkollama list montre votre nouveau modèle aux côtés des autres. La taille affichée est identique à celle du modèle de base — c'est une référence, pas une copie.
- 04Testollama run mon-modele. Le system prompt et les paramètres sont déjà appliqués. Vous pouvez aussi pointer une interface (Open WebUI, LM Studio) sur cette variante via l'API.
#3. Example: concise French assistant
Goal: an assistant that consistently responds in French, without preambles such as "Of course, here’s…", without apologetic formulas, and with calibrated length.
Creation and testing:
A repeat_penalty of 1.15 (instead of the default 1.1) stops models from repeating themselves. A num_ctx of 8192 lets you paste somewhat longer documents without breaking the default 2048-token context.
#4. Example: quiet Python coding
Goal: a coding companion that returns Python ready to paste, with minimal surrounding prose. Ideal for workflows where you want to pipe the output into an editor or script.
A few notable technical choices:
- FROM qwen3-coder:30b
- A code-specialized model from 2026 (MoE 30B-A3B, only 3B active parameters, 256k context, ~19 GB in Q4). On a 24 GB card (RTX 4090) or a 32 GB Mac, it runs at 50–80 tokens/s thanks to its 3B active parameters. If you only have 16 GB of VRAM, replace it with gpt-oss:20b (14 GB) or devstral:24b.
- temperature 0.2
- Low for coding. At 0.7 (default), the model improvises variable names and sometimes invents APIs. For code, you want determinism.
- num_ctx 16384
- Extended window for analyzing entire files. Warning: it multiplies the VRAM consumed by the context.
- stop
- The model stops after the first code block. This avoids post-code explanations cluttering the output.
#5. Example: strict EN→FR translator
Goal: a translator that returns ONLY the translation, without quotation marks, without "Translation:", and without stylistic variation. The typical use case: integration into an n8n pipeline or a bash script where the output will be consumed raw.
Adding MESSAGE to anchor the format through few-shot prompting makes the behavior even more predictable:
Pipeline usage:
#6. Managing custom models
Over time, you'll accumulate several variants. A few useful commands to stay in control:
#7. Which base model should you choose?
Every Modelfile starts with FROM: choosing the base model determines 90% of the final result. A perfect system prompt can never make up for a base that is poorly suited to your use case or too large for your VRAM.
- General-purpose assistant
- A recent 8B to 14B model (from the Qwen or Gemma family) in Q4: responsive, good in French, and leaves room for context.
- Code
- A code-specialized model—MoE models such as Qwen3-Coder 30B-A3B are very fast with 20 GB of memory or more; below that, a 7–9B code model remains a good choice.
- Formal French / writing
- Models Mistral (Nemo, Magistral) retain the advantage in French at the same size.
- Small configuration (8 GB of VRAM)
- Stick to 4B to 9B in Q4: a model that spills into RAM ruins responsiveness, regardless of the Modelfile.
To compare models by VRAM, license, and use case, the QuelLLM catalog filters models compatible with your machine—each entry provides the ollama run exacte command to put in your FROM.
#Troubleshooting
- The model ignores the system prompt
- Make sure you're actually launching your variant (ollama run mon-modele) and not the base model. Confirm with ollama show --system mon-modele.
- Truncated responses
- Increase num_predict (default 128 in some configs) to 2048 or higher. Or add PARAMETER num_predict -1 to disable the limit.
- The model does not follow a strict rule
- Restate the rule in uppercase, add "ABSOLUTELY" or "NEVER", and add an example via MESSAGE. Small models (under ~10B) struggle with isolated negations.
- Error "Error: invalid model reference"
- The FROM model is not installed locally. Run ollama pull <modele> before the ollama create.
- VRAM saturated after creation
- A high num_ctx multiplies consumption. If you go from 2048 to 16384, expect +1 to +3 GB of VRAM depending on the model size. Drop back to 8192 if memory gets tight.
#Go further
The Modelfile is the basic building block for specializing Ollama without touching the weights. Three natural directions to explore next:
- Mastering system prompts in greater depth
- The guide to system prompts explains how to structure roles, constraints, and examples for truly reliable behavior—useful as soon as you exceed 5 lines of SYSTEM.
- Adjust temperature, top-p, and top-k
- Before touching the model, knowing precisely what each sampling parameter changes avoids a lot of back-and-forth. The dedicated guide covers all useful settings.
- Expose your variants through Open WebUI
- Once you create your custom models, Open WebUI lists them automatically. You can switch between assistant-fr, coder-py, and translator-en-fr with one click from the interface.
How do you list the installed Ollama models?+
Where are the Ollama models stored?+
How do you delete a model to free up space?+
Does a custom model take up twice as much disk space?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.