Intermediate 11 minStack

LlamaIndex in pratique

Direct response

LlamaIndex is a Python framework that manages the entire RAG pipeline: document loading, chunking, embeddings, indexing, and queries with sources. Locally, it connects to Ollama and embeddings such as BGE-M3, but its defaults call OpenAI: define Settings.llm and Settings.embed_model before indexing, and configure context_window and request_timeout.

LlamaIndex reduces a RAG pipeline to a few lines, but its defaults (OpenAI, context window, 30-second timeout) trip up local installations, and older agent tutorials no longer work. You’ll learn how to build a fully local pipeline, choose a query mode, add a reranker, and write an agent with the current API.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#LlamaIndex in practice: what it does and when to adopt it

LlamaIndex is a Python framework that handles the entire RAG pipeline: loading documents, splitting them into chunks (nodes), converting them into vectors, indexing them, and then querying the index with an LLM that cites its sources. Version 0.14.25, published on PyPI on September 21, 2026, requires Python 3.10 or later. A minimal RAG fits in about ten lines; the framework becomes useful when you need to vary the sources, change the response mode, add a reranker, or connect an agent. Locally, it works with Ollama for the LLM and an embeddings model of your choice, provided you disable its default settings, which call OpenAI. This guide builds a fully local RAG, shows the settings that matter, and fixes several outdated examples still found online.

Clear abstractions
Document, Node, VectorStoreIndex, retriever, query engine: every RAG step has a dedicated, replaceable object.
Data connectors
SimpleDirectoryReader reads PDF, Word, PowerPoint, Markdown, image, and audio files; other readers support Notion, Google Docs, Slack, and Discord.
Request modes
Several synthesis strategies and more sophisticated engines (sub-questions, routing between indexes) to go beyond simple vector search.
Local is possible
Ollama and Hugging Face embeddings or Ollama connect through dedicated integration packages.

#LlamaIndex RAG steps and their default settings

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before writing code, understand the pipeline and, above all, its defaults: that's where most surprises in local setups hide.

LlamaIndex pipeline: objects and defaults
StepObject or settingA default value worth knowing
BreakdownSentenceSplitter, Settings.chunk_size1,024-token chunk size, 20-token overlap
EmbeddingsSettings.embed_modelOpenAI's text-embedding-ada-002, according to the documentation
LLMSettings.llmOpenAI's gpt-3.5-turbo, according to the getting-started tutorial
StorageVectorStoreIndex, StorageContextIn memory; must be explicitly persisted to disk
Summaryresponse_modecompact: concatenates as many chunks as the window allows
LLM Ollamarequest_timeout30 seconds by default, often too short locally
!
Without configuration, LlamaIndex calls OpenAI
The documentation specifies that LlamaIndex uses the OpenAI API by default for the LLM and embeddings. With an OPENAI_API_KEY present in the environment, your documents would therefore be sent to OpenAI for vectorization, without any warning. For confidential RAG, always define Settings.llm and Settings.embed_model before indexing, as shown in the next section.

#Installation for a 100% local RAG

The pip install llama-index command installs a starter bundle containing llama-index-core, OpenAI integrations for the LLM and embeddings, and file readers. For local use, add the Ollama and embedding integrations (llama-index-embeddings-ollama if you prefer OllamaEmbedding); the bundle's OpenAI packages remain installed but unused as long as you configure Settings.

Terminal
pip install llama-index \
            llama-index-llms-ollama \
            llama-index-embeddings-huggingface

ollama pull mistral

A package name gives you the import: llama-index-llms-ollama corresponds to llama_index.llms.ollama. For memory, count the LLM weights (about 5 GB for a 7-8B model in Q4, according to the site's rule of thumb), the embedding model, and the context: a 16 GB workstation is suitable for a small corpus; more gives you headroom.

#A RAG in ten lines (with its default OpenAI settings)

Python
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader

docs = SimpleDirectoryReader("./docs").load_data()
index = VectorStoreIndex.from_documents(docs)

query = index.as_query_engine()
print(query.query("Résume les points clés du contrat X"))

This code is valid, but it uses OpenAI’s default models: it fails without a key, and with a key it sends your text to the provider. It serves as a skeleton. The next section adds the few lines that make it local.

#Go local with Ollama and French embeddings

Three configuration blocks are enough: the LLM, the embeddings model, and the chunking. For French, BGE-M3 is a common choice: its model card lists more than 100 languages and inputs up to 8,192 tokens. Downloading the model takes about 2.3 GB (the pytorch_model.bin file from the Hugging Face repository), and it happens only once.

Python
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

# Configuration globale : tout est local
Settings.llm = Ollama(
    model="mistral",
    base_url="http://localhost:11434",
    request_timeout=120.0,   # le défaut est de 30 s
    context_window=8000,     # transmis à Ollama comme num_ctx
)
Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-m3")
Settings.chunk_size = 700
Settings.chunk_overlap = 100

# Pipeline
docs = SimpleDirectoryReader("./docs").load_data()
index = VectorStoreIndex.from_documents(docs, show_progress=True)

# Persister sur disque
index.storage_context.persist(persist_dir="./storage")

# Requêter
query_engine = index.as_query_engine(similarity_top_k=5)
reponse = query_engine.query("Quels sont les risques identifiés ?")
print(reponse)
for src in reponse.source_nodes:
    print(f"  - {src.metadata.get('file_name')} ({src.score:.2f})")
Python
from llama_index.core import load_index_from_storage, StorageContext

storage = StorageContext.from_defaults(persist_dir="./storage")
index = load_index_from_storage(storage)

#Why context_window and request_timeout matter

The source code for the Ollama integration in LlamaIndex passes the context_window value to Ollama under the name num_ctx. Without this, Ollama applies its default window of about 4,000 tokens with less than 24 GiB of VRAM, according to its documentation. The math is quick: five 700-token passages amount to 3,500 tokens, before the question, the prompt-template instructions, and the response. With 4,000 tokens, the prompt overflows and the context is truncated, often without any visible error. With 8,000, there is headroom.

Timeout is the other trap: the Ollama client in LlamaIndex defaults to 30 seconds. An initial call that loads the model into memory and processes a long prompt may exceed it. The 120 seconds in the example is a starting point to adjust for your hardware.

→
Embeddings via Ollama instead of Hugging Face
If you don't want to install PyTorch for embeddings, LlamaIndex offers OllamaEmbedding: it uses the already-running Ollama server, with an embedding model pulled by ollama pull. The Ollama library offers bge-m3. Changing the embedding model requires reindexing all documents because vectors from two models are not comparable.

#Loading your files: SimpleDirectoryReader settings

SimpleDirectoryReader reads an entire folder and handles many formats: PDF, Word, PowerPoint, Markdown, images, audio, and video. A few parameters prevent you from indexing just anything.

Python
from llama_index.core import SimpleDirectoryReader

docs = SimpleDirectoryReader(
    input_dir="./docs",
    required_exts=[".pdf", ".docx"],  # ne charger que ces formats
    num_files_limit=100,              # plafond pour un premier essai
).load_data()

Connectors to services (Notion, Google Docs, Slack, Discord) fetch data online: this is unavoidable because the data resides there. Once the documents are loaded, indexing and querying remain local, with your embeddings and your LLM Ollama. Start with a few files: a scanned PDF without a text layer will produce nothing until it has gone through OCR, as described in the Tesseract guide.

#Choose a query mode and the cost in LLM calls

The synthesis mode determines how many times the LLM is called, and therefore the local response time. The table covers the documented modes and their cost, with an example using five passes of 700 tokens and an 8,000-token window (estimate).

Response modes and number of LLM calls
ModePrinciple (documentation)Calls in the example
compact (default)Concatenates as many chunks as the window allows, then queries1 call: 3,500 tokens fit into 8,000
refineProcesses the chunks one by one, with one call per chunk5 calls in a row
tree_summarizeQueries in batches, then recursively summarizes the responses1 call if everything fits in the window; otherwise, several calls, followed by a final summary
simple_summarizeTruncates everything to fit in a single prompt1 call, with loss of detail
no_textOnly runs the retriever, without calling the LLM0 calls; useful for debugging the search

For a factual question, keep it compact. To summarize a long document, tree_summarize is designed for the job, at the cost of multiple calls: on a local machine, expect it to take several times as long as a simple response. no_text is valuable for checking what the search returns without waiting for the LLM.

#Add a local reranker

When the right answer is in the top 20 passages but not the top 5, a reranker reorders the candidates before they are sent to the LLM. LlamaIndex documentation recommends SentenceTransformerRerank as the default when running locally without an API key, a cross-encoder via sentence-transformers, and cites Qwen3-Reranker-0.6B for better multilingual quality.

Python
from llama_index.core.postprocessor import SentenceTransformerRerank

reranker = SentenceTransformerRerank(
    model="cross-encoder/ms-marco-MiniLM-L2-v2",
    top_n=3,
)
query_engine = index.as_query_engine(
    similarity_top_k=15,             # large pour rattraper la bonne réponse
    node_postprocessors=[reranker],  # puis resserre à 3
)

The model used here is the one from the official example, chosen for its speed; it is designed for English. For French documents, test a multilingual reranker. The dedicated guide details the selection and evaluation.

#Subquestions and routing

SubQuestionQueryEngine
Breaks a complex question into subquestions sent to query tools, then synthesizes the results. “Compare the 2024 and 2025 strategies” becomes two separate searches.
RouterQueryEngine
Chooses the right engine for each question from several options (for example, a summary index or a vector index).

These two engines multiply LLM calls: decomposing a question into three subquestions adds the generation of the subquestions, three answers, and the final synthesis, for at least five calls. On local hardware, reserve them for questions that warrant it.

#Agents: the current API is no longer the one used in older tutorials

Many tutorials use ReActAgent.from_tools. This class no longer exists in the current source: the agent/react/base.py module that contained it has disappeared from the repository. Agents are now asynchronous workflows: FunctionAgent (an agent that calls functions or tools), the workflow version of ReActAgent, and AgentWorkflow for orchestrating multiple agents. The official local tutorial builds the agent with AgentWorkflow.from_tools_or_functions and runs it with await agent.run.

Python
import asyncio
from llama_index.core import Settings
from llama_index.core.agent.workflow import AgentWorkflow

async def search_documents(query: str) -> str:
    """Répond aux questions sur les contrats clients."""
    response = await query_engine.aquery(query)
    return str(response)

agent = AgentWorkflow.from_tools_or_functions(
    [search_documents],
    llm=Settings.llm,
    system_prompt="Tu réponds uniquement à partir des contrats indexés.",
)

async def main():
    reponse = await agent.run("Y a-t-il une clause de non-concurrence chez Acme Corp ?")
    print(str(reponse))

asyncio.run(main())

The function's name, description, and arguments (its docstring) are passed to the LLM, which decides whether to call it, so make this description precise. FunctionAgent relies on native function calls; with a local model, choose one that supports tools in Ollama, and keep the RAG simple if yours cannot.

#LlamaIndex, LangChain, or custom RAG: how to choose

Three approaches to local RAG
CriterionCustom RAG (ChromaDB, embeddings)LlamaIndexLangChain
Main objectiveUnderstand every step, total controlDocument pipeline with ready-made componentsTool and agent orchestration
Startup timeLonger: everything still needs to be writtenShort for a first RAGShort for a chain, longer for a full RAG
CustomizationUnlimited, at your expensePer-component settings, replaceable objectsVery flexible, more verbose
Main riskReinventing an existing pipelineOpenAI defaults, rapidly evolving APIFast-evolving API

Rule of thumb: a document RAG with cited sources, varied formats, and a few query options can be built quickly with LlamaIndex. If you first want to understand the mechanism, write the homemade version once: the ChromaDB guide shows the steps. For an agent that controls many external tools, compare LangChain or LlamaIndex workflows based on your habits.

#The pitfalls that cost you time

Reindex on every run
Without persist, the index stays in memory and disappears when the script closes. Persist it, then reload it with load_index_from_storage.
Switch models without reindexing
Vectors from two embedding models are not comparable: changing embed_model or chunking requires rebuilding the index.
Context silently truncated
A prompt longer than the Ollama window is truncated: check context_window if responses ignore the last passages.
Copy an old tutorial
Agent APIs have changed: verify that the imports in an example exist in the installed version (0.14.25 at the time of writing).
Never measure
An unevaluated RAG system drifts without you knowing it: the Ragas guide shows you how to quantify it.
FAQ
Does LlamaIndex run entirely locally?+
Yes, provided you define Settings.llm (for example, Ollama) and Settings.embed_model (Hugging Face or Ollama) before indexing. Without this, the documentation says LlamaIndex uses OpenAI models by default. Only connectors to online services, such as Notion or Slack, necessarily leave the machine to retrieve data.
Which embeddings model should you choose for French-language documents?+
BGE-M3 is a common choice: its model card lists more than 100 languages, inputs of up to 8,192 tokens, and a download of about 2.3 GB. Benchmarks vary with your documents: test it on your own questions before generalizing, and consult the French embeddings guide to compare alternatives.
Why does my LlamaIndex RAG forget passages?+
Often because the context is truncated: Ollama defaults to roughly 4,000 tokens under 24 GiB of VRAM, and five 700-token passages already use 3,500. Set context_window in the Ollama object, increase the available memory if necessary, or reduce similarity_top_k.
How can you avoid reindexing on every launch?+
Persist the index with index.storage_context.persist(persist_dir="./storage"), then reload it with StorageContext.from_defaults and load_index_from_storage. By default, LlamaIndex keeps data in memory and loses it when the script closes. If you change the embedding model or chunk size, the computed vectors change, so you must rebuild the index and persist it again.
ReActAgent.from_tools no longer works. What should I do?+
This API from older tutorials has disappeared from the current sources. Use workflow agents: AgentWorkflow.from_tools_or_functions or FunctionAgent, with asynchronous functions and await agent.run. Check that your Ollama model supports tools; otherwise, stick with a simple RAG without an agent, or test the ReActAgent workflow, which does not require native function calling.
LlamaIndex or LangChain for local RAG?+
LlamaIndex focuses on data and RAG: loading, indexing, and query modes. LangChain focuses on orchestrating tools and agents. To chat with your documents, LlamaIndex gets you started faster; for a multi-tool agent, compare both on a real use case and keep the one whose API you understand best.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.