LlamaIndex in pratique
LlamaIndex is a Python framework that manages the entire RAG pipeline: document loading, chunking, embeddings, indexing, and queries with sources. Locally, it connects to Ollama and embeddings such as BGE-M3, but its defaults call OpenAI: define Settings.llm and Settings.embed_model before indexing, and configure context_window and request_timeout.
LlamaIndex reduces a RAG pipeline to a few lines, but its defaults (OpenAI, context window, 30-second timeout) trip up local installations, and older agent tutorials no longer work. You’ll learn how to build a fully local pipeline, choose a query mode, add a reranker, and write an agent with the current API.
#LlamaIndex in practice: what it does and when to adopt it
LlamaIndex is a Python framework that handles the entire RAG pipeline: loading documents, splitting them into chunks (nodes), converting them into vectors, indexing them, and then querying the index with an LLM that cites its sources. Version 0.14.25, published on PyPI on September 21, 2026, requires Python 3.10 or later. A minimal RAG fits in about ten lines; the framework becomes useful when you need to vary the sources, change the response mode, add a reranker, or connect an agent. Locally, it works with Ollama for the LLM and an embeddings model of your choice, provided you disable its default settings, which call OpenAI. This guide builds a fully local RAG, shows the settings that matter, and fixes several outdated examples still found online.
- Clear abstractions
- Document, Node, VectorStoreIndex, retriever, query engine: every RAG step has a dedicated, replaceable object.
- Data connectors
- SimpleDirectoryReader reads PDF, Word, PowerPoint, Markdown, image, and audio files; other readers support Notion, Google Docs, Slack, and Discord.
- Request modes
- Several synthesis strategies and more sophisticated engines (sub-questions, routing between indexes) to go beyond simple vector search.
- Local is possible
- Ollama and Hugging Face embeddings or Ollama connect through dedicated integration packages.
#LlamaIndex RAG steps and their default settings
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
Before writing code, understand the pipeline and, above all, its defaults: that's where most surprises in local setups hide.
| Step | Object or setting | A default value worth knowing |
|---|---|---|
| Breakdown | SentenceSplitter, Settings.chunk_size | 1,024-token chunk size, 20-token overlap |
| Embeddings | Settings.embed_model | OpenAI's text-embedding-ada-002, according to the documentation |
| LLM | Settings.llm | OpenAI's gpt-3.5-turbo, according to the getting-started tutorial |
| Storage | VectorStoreIndex, StorageContext | In memory; must be explicitly persisted to disk |
| Summary | response_mode | compact: concatenates as many chunks as the window allows |
| LLM Ollama | request_timeout | 30 seconds by default, often too short locally |
#Installation for a 100% local RAG
The pip install llama-index command installs a starter bundle containing llama-index-core, OpenAI integrations for the LLM and embeddings, and file readers. For local use, add the Ollama and embedding integrations (llama-index-embeddings-ollama if you prefer OllamaEmbedding); the bundle's OpenAI packages remain installed but unused as long as you configure Settings.
A package name gives you the import: llama-index-llms-ollama corresponds to llama_index.llms.ollama. For memory, count the LLM weights (about 5 GB for a 7-8B model in Q4, according to the site's rule of thumb), the embedding model, and the context: a 16 GB workstation is suitable for a small corpus; more gives you headroom.
#A RAG in ten lines (with its default OpenAI settings)
This code is valid, but it uses OpenAI’s default models: it fails without a key, and with a key it sends your text to the provider. It serves as a skeleton. The next section adds the few lines that make it local.
#Go local with Ollama and French embeddings
Three configuration blocks are enough: the LLM, the embeddings model, and the chunking. For French, BGE-M3 is a common choice: its model card lists more than 100 languages and inputs up to 8,192 tokens. Downloading the model takes about 2.3 GB (the pytorch_model.bin file from the Hugging Face repository), and it happens only once.
#Why context_window and request_timeout matter
The source code for the Ollama integration in LlamaIndex passes the context_window value to Ollama under the name num_ctx. Without this, Ollama applies its default window of about 4,000 tokens with less than 24 GiB of VRAM, according to its documentation. The math is quick: five 700-token passages amount to 3,500 tokens, before the question, the prompt-template instructions, and the response. With 4,000 tokens, the prompt overflows and the context is truncated, often without any visible error. With 8,000, there is headroom.
Timeout is the other trap: the Ollama client in LlamaIndex defaults to 30 seconds. An initial call that loads the model into memory and processes a long prompt may exceed it. The 120 seconds in the example is a starting point to adjust for your hardware.
#Loading your files: SimpleDirectoryReader settings
SimpleDirectoryReader reads an entire folder and handles many formats: PDF, Word, PowerPoint, Markdown, images, audio, and video. A few parameters prevent you from indexing just anything.
Connectors to services (Notion, Google Docs, Slack, Discord) fetch data online: this is unavoidable because the data resides there. Once the documents are loaded, indexing and querying remain local, with your embeddings and your LLM Ollama. Start with a few files: a scanned PDF without a text layer will produce nothing until it has gone through OCR, as described in the Tesseract guide.
#Choose a query mode and the cost in LLM calls
The synthesis mode determines how many times the LLM is called, and therefore the local response time. The table covers the documented modes and their cost, with an example using five passes of 700 tokens and an 8,000-token window (estimate).
| Mode | Principle (documentation) | Calls in the example |
|---|---|---|
| compact (default) | Concatenates as many chunks as the window allows, then queries | 1 call: 3,500 tokens fit into 8,000 |
| refine | Processes the chunks one by one, with one call per chunk | 5 calls in a row |
| tree_summarize | Queries in batches, then recursively summarizes the responses | 1 call if everything fits in the window; otherwise, several calls, followed by a final summary |
| simple_summarize | Truncates everything to fit in a single prompt | 1 call, with loss of detail |
| no_text | Only runs the retriever, without calling the LLM | 0 calls; useful for debugging the search |
For a factual question, keep it compact. To summarize a long document, tree_summarize is designed for the job, at the cost of multiple calls: on a local machine, expect it to take several times as long as a simple response. no_text is valuable for checking what the search returns without waiting for the LLM.
#Add a local reranker
When the right answer is in the top 20 passages but not the top 5, a reranker reorders the candidates before they are sent to the LLM. LlamaIndex documentation recommends SentenceTransformerRerank as the default when running locally without an API key, a cross-encoder via sentence-transformers, and cites Qwen3-Reranker-0.6B for better multilingual quality.
The model used here is the one from the official example, chosen for its speed; it is designed for English. For French documents, test a multilingual reranker. The dedicated guide details the selection and evaluation.
#Subquestions and routing
- SubQuestionQueryEngine
- Breaks a complex question into subquestions sent to query tools, then synthesizes the results. “Compare the 2024 and 2025 strategies” becomes two separate searches.
- RouterQueryEngine
- Chooses the right engine for each question from several options (for example, a summary index or a vector index).
These two engines multiply LLM calls: decomposing a question into three subquestions adds the generation of the subquestions, three answers, and the final synthesis, for at least five calls. On local hardware, reserve them for questions that warrant it.
#Agents: the current API is no longer the one used in older tutorials
Many tutorials use ReActAgent.from_tools. This class no longer exists in the current source: the agent/react/base.py module that contained it has disappeared from the repository. Agents are now asynchronous workflows: FunctionAgent (an agent that calls functions or tools), the workflow version of ReActAgent, and AgentWorkflow for orchestrating multiple agents. The official local tutorial builds the agent with AgentWorkflow.from_tools_or_functions and runs it with await agent.run.
The function's name, description, and arguments (its docstring) are passed to the LLM, which decides whether to call it, so make this description precise. FunctionAgent relies on native function calls; with a local model, choose one that supports tools in Ollama, and keep the RAG simple if yours cannot.
#LlamaIndex, LangChain, or custom RAG: how to choose
| Criterion | Custom RAG (ChromaDB, embeddings) | LlamaIndex | LangChain |
|---|---|---|---|
| Main objective | Understand every step, total control | Document pipeline with ready-made components | Tool and agent orchestration |
| Startup time | Longer: everything still needs to be written | Short for a first RAG | Short for a chain, longer for a full RAG |
| Customization | Unlimited, at your expense | Per-component settings, replaceable objects | Very flexible, more verbose |
| Main risk | Reinventing an existing pipeline | OpenAI defaults, rapidly evolving API | Fast-evolving API |
Rule of thumb: a document RAG with cited sources, varied formats, and a few query options can be built quickly with LlamaIndex. If you first want to understand the mechanism, write the homemade version once: the ChromaDB guide shows the steps. For an agent that controls many external tools, compare LangChain or LlamaIndex workflows based on your habits.
#The pitfalls that cost you time
- Reindex on every run
- Without persist, the index stays in memory and disappears when the script closes. Persist it, then reload it with load_index_from_storage.
- Switch models without reindexing
- Vectors from two embedding models are not comparable: changing embed_model or chunking requires rebuilding the index.
- Context silently truncated
- A prompt longer than the Ollama window is truncated: check context_window if responses ignore the last passages.
- Copy an old tutorial
- Agent APIs have changed: verify that the imports in an example exist in the installed version (0.14.25 at the time of writing).
- Never measure
- An unevaluated RAG system drifts without you knowing it: the Ragas guide shows you how to quantify it.
- Local RAG with ChromaDB and Ollama: Python tutorial
- Add a reranker to your pipeline
- The best French embedding models
- Chunking strategies
- Ragas: evaluate your local RAG with numbers
- Build a local AI agent in Python with LangChain and Ollama
- Source: getting-started tutorial with local models
- Source: LlamaIndex installation
- Source: response-synthesis modes
- Source: index persistence and reloading
- Source: BGE-M3 model sheet
Does LlamaIndex run entirely locally?+
Which embeddings model should you choose for French-language documents?+
Why does my LlamaIndex RAG forget passages?+
How can you avoid reindexing on every launch?+
ReActAgent.from_tools no longer works. What should I do?+
LlamaIndex or LangChain for local RAG?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.