Your first conversation locale
To query a local LLM, launch a model with ollama run, enter a complete request in one or two sentences with the text to process, read the response, then correct it with a precise instruction instead of starting over. Nothing you write leaves your machine. In about fifteen minutes, you will have created a summary, an email, a little code, and learned how to fix a failed response.
Ollama is installed, and a blinking cursor is waiting in front of you. This first-day tutorial guides you through a real conversation: launch a model suited to your memory, write prompts that work, use session commands, spot when the model goes off track, and avoid context and memory pitfalls.
#What you'll be able to do in this tutorial
You will ask for a summary, have an email written, request a small script, test a question outside the model's scope, then correct an answer that is heading in the wrong direction. These are the actions that make you self-sufficient. The guide assumes Ollama is already installed; if not, start with your system's installation guide. To go further on how to phrase an instruction, the guide to prompting basics takes over.
#Choose a model that fits on your machine
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The model is a file several gigabytes in size: it must fit in your card's video memory, or in the RAM of a Mac or PC without a dedicated card. The Gemma 4 page in the Ollama library lists the variants and their download sizes. A variant that is too large for your machine will work, but slowly, spilling over onto the processor.
| Label | Download size | Memory requirements |
|---|---|---|
| gemma4:e2b | 4.6 to 7.5 GB | 8 GB machines |
| gemma4 (e4b, default) | 6.6 to 9.5 GB | 12 to 16 GB machines |
| gemma4:12b | 7.7 to 8.0 GB | Machines with 16 GB or more |
| gemma4:26b | 16 to 19 GB | 24 GB or larger cards |
| gemma4:31b | 19 to 20 GB | 24 to 32 GB cards, reduced context |
The right-hand column is our cautious estimate: add the context memory and a margin for the system to the sizes above. When in doubt, choose the smallest variant: a fast response from a small model teaches you more than a very slow response from a large one. You can switch models at any time.
#Your first session, step by step
Allow about fifteen minutes to go through the eight steps below. They cover everything from launching to verifying a response. Each step is independent: you can stop after the second and come back later. The only thing you lose by closing the session is the conversation history, which the terminal does not retain.
#1. Launch the model
The ollama run gemma4 command appears in the Ollama documentation as an example of running a model. On the first launch, the model is downloaded: the duration depends on its size and your connection. Then a prompt appears, and the model waits for your message. Nothing happens until you type.
#2. Talk to it in complete sentences
First classic mistake: writing keywords as you would for a search engine. A language model does not search for documents; it continues text. So describe what you want, including the context and expected result.
A structure that works almost every time: the task, the data to process, and the expected format. Four exercises to practice: ask for a summary of personal notes, draft a three-line cancellation email, ask for a Python function that counts the words in a file, and ask a question the model can’t answer, such as the contents of a document you haven’t given it.
#3. Test what the model does not know
The final exercise is the most instructive. A model does not access the internet or your files unless a tool is connected: it knows only its training and the conversation text. Ask it for the contents of an imaginary document or about a recent event, and observe: a good model says it does not know; a less reliable model invents a plausible answer. This behavior is called a hallucination. The resulting rule: the more precise, recent, or specific the information is to you, the more you must provide it in the prompt or verify it. Our guide to hallucinations details the countermeasures.
#4. Fix it without starting over
Is the model responding beside the point? Don't just type “no, try again”: it has not understood anything more. Add a precise constraint that says what is wrong and what you expect.
- Length
- Number of chips, lines, or words.
- Ton
- Formal, casual, journalistic, instructional.
- Form
- List, table, JSON, email, code.
- Prohibitions phrased positively
- Prefer “start directly with the answer” to “don’t say here’s”.
#5. Session commands
In the session, lines that begin with a slash are not sent to the model: they control Ollama. The /? command displays the complete list.
- /set parameter num_ctx 8192
- Sets the context size for the session: Ollama’s FAQ cites this command to change the window.
- /set system
- Defines a system message, a permanent role for the session.
- /show info
- Displays information about the loaded model: architecture, size, context length, and quantization.
- /clear
- Clear the current conversation's context without unloading the model.
- /bye
- Close the session, or press Ctrl+D.
Once the session is closed, the model remains in memory for five minutes by default, according to the Ollama FAQ, which speeds up a new launch. To unload it immediately, run ollama stop followed by the model name; ollama ps lists what is loaded and shows the portion running on the GPU and CPU.
#6. Hold a multi-turn conversation
The model only remembers what is in the context window: on every turn, the application sends it the entire history. So you can chain requests that build on one another.
The limitation is context size. According to Ollama's documentation, the default is around 4,000 tokens below 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above that. When the history exceeds the window, the oldest exchanges are lost. If the model “forgets” the beginning of a long session, increase the window with the num_ctx command, keeping in mind that each context token consumes memory.
#7. Ask for code, and verify it
Code is a good testing ground because the result can be tested. Give it the language, input, expected output, and one constraint: “Write a Python function that takes the path to a text file and returns the number of words. Respond with the code only, followed by one explanatory sentence.” Copy the code into a file, run it on a case whose answer you know, and if an error appears, paste the complete error message into the conversation and ask it to fix it. This request, execution, correction cycle works better than demanding perfect code on the first try.
#8. Read a response critically
A fluent response isn’t an accurate response. The model produces plausible text; it doesn’t verify anything. Make a habit of checking anything that matters before using it.
- Numbers and dates
- Cross-check them against a source: these are the elements models most easily make up.
- Names, references, and citations
- Verify that they exist. A model can fabricate a book title or statute that sounds plausible.
- Code
- Run it before relying on it, and review what it does to your files.
- Consistency
- Ask the same question in two ways: if the answers differ, be wary of both.
A good practice is to provide the reference text yourself and ask the model to cite the passage that supports each claim. This lets you verify them in a few seconds instead of searching for everything yourself.
#Stay in the terminal or switch to an interface?
| Path | Advantage | Limit |
|---|---|---|
| Terminal (ollama run) | No additional installation, ideal for a first try | No organized history, little comfort for long texts |
| Chat interface (Open WebUI, LM Studio, Jan) | History, attachments, readable layout | An additional installation |
| API (scripts) | Automation, integration with your tools | Asks for a little code |
Start with the terminal to understand the principle, then switch to an interface as soon as you paste long texts or want to retrieve your conversations. The two work together: the model downloaded once with Ollama also serves the interfaces, with no second download, as long as they use the same local server.
#When something goes wrong
| Symptom | Likely cause | Solution |
|---|---|---|
| Very slow responses | The model spills over onto the processor | Run ollama ps, then choose a smaller variant or reduce the context |
| The model forgets the beginning | Full context window | Increase num_ctx or start over with a summary |
| English response | English-oriented prompt or model | Write « réponds en français » in the request or in /set system |
| Invented text | Information missing from the prompt | Provide the source text and ask it to quote the passage |
| Model not found error | Incorrect name or label | Check the spelling on the model’s page in the Ollama library |
For finer control over creativity, the Modelfile documentation for Ollama summarizes the effect of temperature: the higher it is, the more creative the model; the lower it is, the more consistent it is. The guide to temperature and top-p details values to try.
#And after that first conversation?
- A permanent role
- The guide to system prompts shows how to give the model stable rules.
- A more comfortable interface
- Open WebUI or LM Studio add history, file attachments, and a readable layout.
- Your own documents
- RAG lets you query your files without pasting them in every time.
How do you query an LLM for the first time?+
Do you need an internet connection to chat with a local model?+
How do you exit a Ollama conversation?+
Why does the model forget the beginning of the conversation?+
Which model should you choose to get started?+
Can the model read my files or the internet?+
#Go further
- Install Ollama in 5 minutes
- Prompting basics
- Master system prompts
- Understanding the context window
- Hallucinations: how to limit them
- Open WebUI with Ollama: complete guide
- Source: Ollama FAQ
- Source: Ollama, context length
- Source: Ollama library, Gemma 4
- Source: Ollama, Modelfile reference
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.