Beginner 10 minPrompting

Your first conversation locale

Direct response

To query a local LLM, launch a model with ollama run, enter a complete request in one or two sentences with the text to process, read the response, then correct it with a precise instruction instead of starting over. Nothing you write leaves your machine. In about fifteen minutes, you will have created a summary, an email, a little code, and learned how to fix a failed response.

Ollama is installed, and a blinking cursor is waiting in front of you. This first-day tutorial guides you through a real conversation: launch a model suited to your memory, write prompts that work, use session commands, spot when the model goes off track, and avoid context and memory pitfalls.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What you'll be able to do in this tutorial

You will ask for a summary, have an email written, request a small script, test a question outside the model's scope, then correct an answer that is heading in the wrong direction. These are the actions that make you self-sufficient. The guide assumes Ollama is already installed; if not, start with your system's installation guide. To go further on how to phrase an instruction, the guide to prompting basics takes over.

#Choose a model that fits on your machine

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The model is a file several gigabytes in size: it must fit in your card's video memory, or in the RAM of a Mac or PC without a dedicated card. The Gemma 4 page in the Ollama library lists the variants and their download sizes. A variant that is too large for your machine will work, but slowly, spilling over onto the processor.

Gemma variants 4 in the Ollama library (size indicated by Ollama)
LabelDownload sizeMemory requirements
gemma4:e2b4.6 to 7.5 GB8 GB machines
gemma4 (e4b, default)6.6 to 9.5 GB12 to 16 GB machines
gemma4:12b7.7 to 8.0 GBMachines with 16 GB or more
gemma4:26b16 to 19 GB24 GB or larger cards
gemma4:31b19 to 20 GB24 to 32 GB cards, reduced context

The right-hand column is our cautious estimate: add the context memory and a margin for the system to the sizes above. When in doubt, choose the smallest variant: a fast response from a small model teaches you more than a very slow response from a large one. You can switch models at any time.

#Your first session, step by step

Allow about fifteen minutes to go through the eight steps below. They cover everything from launching to verifying a response. Each step is independent: you can stop after the second and come back later. The only thing you lose by closing the session is the conversation history, which the terminal does not retain.

#1. Launch the model

First launch
ollama run gemma4:e2b

The ollama run gemma4 command appears in the Ollama documentation as an example of running a model. On the first launch, the model is downloaded: the duration depends on its size and your connection. Then a prompt appears, and the model waits for your message. Nothing happens until you type.

i
Your messages stay with you
The Ollama FAQ states that when you run a model locally, Ollama sees neither your prompts nor your data. This is not true for cloud models, which send your messages to the provider’s servers: make sure the model name does not end with cloud.

#2. Talk to it in complete sentences

First classic mistake: writing keywords as you would for a search engine. A language model does not search for documents; it continues text. So describe what you want, including the context and expected result.

An unclear request, then a complete request
# Trop vague
résumer réunion points importants

# Précis
Voici des notes de réunion. Résume-les en 5 puces : d'abord les décisions
prises, puis les actions à faire avec leur responsable.

Notes :
- Alice propose de migrer vers Postgres d'ici juin
- Bob rappelle qu'il faut finir l'audit sécurité avant
- Décision : on migre, mais l'audit passe avant (fin mai)
- Charles s'occupe de l'audit
- Alice s'occupe de la migration

A structure that works almost every time: the task, the data to process, and the expected format. Four exercises to practice: ask for a summary of personal notes, draft a three-line cancellation email, ask for a Python function that counts the words in a file, and ask a question the model can’t answer, such as the contents of a document you haven’t given it.

#3. Test what the model does not know

The final exercise is the most instructive. A model does not access the internet or your files unless a tool is connected: it knows only its training and the conversation text. Ask it for the contents of an imaginary document or about a recent event, and observe: a good model says it does not know; a less reliable model invents a plausible answer. This behavior is called a hallucination. The resulting rule: the more precise, recent, or specific the information is to you, the more you must provide it in the prompt or verify it. Our guide to hallucinations details the countermeasures.

#4. Fix it without starting over

Is the model responding beside the point? Don't just type “no, try again”: it has not understood anything more. Add a precise constraint that says what is wrong and what you expect.

A targeted restart
Refais, mais en 3 puces maximum et sans jargon technique.
Chaque puce doit tenir sur une ligne.
Length
Number of chips, lines, or words.
Ton
Formal, casual, journalistic, instructional.
Form
List, table, JSON, email, code.
Prohibitions phrased positively
Prefer “start directly with the answer” to “don’t say here’s”.

#5. Session commands

In the session, lines that begin with a slash are not sent to the model: they control Ollama. The /? command displays the complete list.

/set parameter num_ctx 8192
Sets the context size for the session: Ollama’s FAQ cites this command to change the window.
/set system
Defines a system message, a permanent role for the session.
/show info
Displays information about the loaded model: architecture, size, context length, and quantization.
/clear
Clear the current conversation's context without unloading the model.
/bye
Close the session, or press Ctrl+D.

Once the session is closed, the model remains in memory for five minutes by default, according to the Ollama FAQ, which speeds up a new launch. To unload it immediately, run ollama stop followed by the model name; ollama ps lists what is loaded and shows the portion running on the GPU and CPU.

#6. Hold a multi-turn conversation

The model only remembers what is in the context window: on every turn, the application sends it the entire history. So you can chain requests that build on one another.

Chain three turns
Je dois écrire un courriel pour annuler une réservation au restaurant.
Fais-le en 3 lignes, poli mais pas guindé.

[réponse du modèle]

Bien. Refais la même, mais plus chaleureuse, et précise que je
reprogrammerai dans 15 jours.

[réponse du modèle]

Parfait. Traduis-la en anglais.

The limitation is context size. According to Ollama's documentation, the default is around 4,000 tokens below 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above that. When the history exceeds the window, the oldest exchanges are lost. If the model “forgets” the beginning of a long session, increase the window with the num_ctx command, keeping in mind that each context token consumes memory.

#7. Ask for code, and verify it

Code is a good testing ground because the result can be tested. Give it the language, input, expected output, and one constraint: “Write a Python function that takes the path to a text file and returns the number of words. Respond with the code only, followed by one explanatory sentence.” Copy the code into a file, run it on a case whose answer you know, and if an error appears, paste the complete error message into the conversation and ask it to fix it. This request, execution, correction cycle works better than demanding perfect code on the first try.

#8. Read a response critically

A fluent response isn’t an accurate response. The model produces plausible text; it doesn’t verify anything. Make a habit of checking anything that matters before using it.

Numbers and dates
Cross-check them against a source: these are the elements models most easily make up.
Names, references, and citations
Verify that they exist. A model can fabricate a book title or statute that sounds plausible.
Code
Run it before relying on it, and review what it does to your files.
Consistency
Ask the same question in two ways: if the answers differ, be wary of both.

A good practice is to provide the reference text yourself and ask the model to cite the passage that supports each claim. This lets you verify them in a few seconds instead of searching for everything yourself.

#Stay in the terminal or switch to an interface?

Three ways to interact with a local model
PathAdvantageLimit
Terminal (ollama run)No additional installation, ideal for a first tryNo organized history, little comfort for long texts
Chat interface (Open WebUI, LM Studio, Jan)History, attachments, readable layoutAn additional installation
API (scripts)Automation, integration with your toolsAsks for a little code

Start with the terminal to understand the principle, then switch to an interface as soon as you paste long texts or want to retrieve your conversations. The two work together: the model downloaded once with Ollama also serves the interfaces, with no second download, as long as they use the same local server.

#When something goes wrong

Common symptoms and solutions
SymptomLikely causeSolution
Very slow responsesThe model spills over onto the processorRun ollama ps, then choose a smaller variant or reduce the context
The model forgets the beginningFull context windowIncrease num_ctx or start over with a summary
English responseEnglish-oriented prompt or modelWrite « réponds en français » in the request or in /set system
Invented textInformation missing from the promptProvide the source text and ask it to quote the passage
Model not found errorIncorrect name or labelCheck the spelling on the model’s page in the Ollama library

For finer control over creativity, the Modelfile documentation for Ollama summarizes the effect of temperature: the higher it is, the more creative the model; the lower it is, the more consistent it is. The guide to temperature and top-p details values to try.

#And after that first conversation?

A permanent role
The guide to system prompts shows how to give the model stable rules.
A more comfortable interface
Open WebUI or LM Studio add history, file attachments, and a readable layout.
Your own documents
RAG lets you query your files without pasting them in every time.
FAQ
How do you query an LLM for the first time?+
Launch a model with ollama run followed by its name, wait for the prompt, then write a complete request: the task, the text to process, and the expected format. Read the response, then correct it with a precise instruction instead of starting over. A few exchanges are enough to understand what the model can do.
Do you need an internet connection to chat with a local model?+
Only to download the model. After that, the conversation runs on your machine, without a connection. The Ollama FAQ states that when you run a model locally, Ollama sees neither your prompts nor your data. Models with the cloud suffix work the other way around on remote servers and require a connection.
How do you exit a Ollama conversation?+
Type /bye or use Ctrl+D to close the session. The model then remains loaded in memory for five minutes by default, speeding up the next launch. To unload it immediately, run ollama stop followed by the model name, and ollama ps to check what is still loaded.
Why does the model forget the beginning of the conversation?+
Because the history exceeds the context window. Ollama uses about 4,000 tokens by default with 24 GiB of VRAM. Increase the value with /set parameter num_ctx in the session, checking that the model still fits on the GPU, or summarize the conversation before continuing.
Which model should you choose to get started?+
Choose the smallest capable model in the family you want to try, suited to your memory: a model that responds quickly teaches you more than a very slow large model. With 8 GB, the e2b variant of Gemma 4 weighs about 4.6 to 7.5 GB. Then move up in size if the quality disappoints you.
Can the model read my files or the internet?+
Not by default. It only knows its training and the conversation text. To have it read your documents or search the web, you need to connect a tool, such as an interface like Open WebUI with RAG or web search. Otherwise, paste the text into the prompt.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.