Advanced 13 minOptimization

Ragas: evaluate your local RAG with chiffres

Direct response

Ragas is an open-source Python library (Apache 2.0 license, more than 15,000 stars on GitHub) that quantifies the quality of a document-retrieval pipeline: faithfulness, answer relevance, context precision, and context recall. It can run entirely with a local judge model and embedding model, avoiding the need to send your documents and answers to a third-party service just to evaluate them.

You tune a document-retrieval pipeline blindly until you have a number: you change chunk sizes, replace the embedding model, and judge by feel on three questions. Ragas is a Python library that puts metrics behind these intuitions—and separates, when an answer is poor, retrieval’s fault from the model’s. It can run entirely with local models, avoiding the need to send your documents to a third-party service for evaluation.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Why “it looks good” isn’t enough

Ragas is an open-source Python library under the Apache 2.0 license that quantifies the quality of a document-retrieval pipeline: the answer's faithfulness to the provided passages, its relevance to the question, and the precision and recall of the retrieved context. This separates retrieval errors from model errors. It can run with a local judge served by Ollama, provided that model follows the structured-output format Ragas requires and its context contains the entire question, passages, and answer. Before adopting it, keep three points in mind: its scores are for comparing two versions of the same system, not for rating absolute quality; the test set matters more than the library; and online tutorials mix the old and new APIs.

When an answer is wrong, two very different causes can hide behind the same symptom: either the retrieved passages did not contain the information, or they did and the model answered beside the point. The fix is not the same—chunking and embeddings in the first case, the model and prompt in the second. Without measurement, you are fixing things at random, and an improvement on three questions can worsen five others without you noticing.

The other pitfall is comparison. “Is the 27-billion-parameter model really better here?” is answered by rerunning the same set of questions, not by discussing it. That is what reproducible evaluation makes possible. Ragas has surpassed 15,000 stars on GitHub. The latest version published on PyPI is 0.4.3, from January 13, 2026, and the repository has changed GitHub organizations (vibrantlabsai). Pin the version you use.

#The four metrics that matter

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
What each measurement diagnoses
MetricQuestion askedWhat it blames when it dropsWhat you need to provide
FaithfulnessIs the response fully supported by the supplied passages? Score: supported claims divided by total claims in the responseThe model makes things up or extrapolatesQuestion, answer, retrieved passages
Response relevanceDoes the answer address the question asked? The judge generates three questions from the answer, then compares their similarity with the original questionThe prompt or model goes off topicQuestion, answer, and an embedding model
Context accuracyAre the relevant passages ranked near the top? Average precision at each rankSearch returns noise or ranks results poorlyQuestion, passages in retrieved order, reference answer
Context reminderDid we retrieve everything we needed? Share of the reference answer's claims supported by the passagesChunks too small, threshold too strict, bad embeddingsQuestion, passages, reference answer

Ragas documentation specifies that answer relevance does not judge correctness: it only measures how well the answer fits the question, and penalizes incomplete answers or ones full of unnecessary details. That is why it does not replace faithfulness. Cross-reading is what makes these figures useful. Low recall with high faithfulness describes an honest but poorly supplied system: you need to improve retrieval. Low faithfulness with good recall describes the opposite: the information was there, but the model embellished. The two are corrected at opposite points in the pipeline.

i
Fidelity is the first metric to monitor
In a professional context, an invented but plausible answer costs more than no answer. A pipeline that admits “the documents don’t say” is preferable to a polished pipeline that occasionally gets things wrong without warning.

#Build a test set

This is the part that requires work, and it can’t be automated without consequences. A useful dataset contains questions users actually asked, including their awkward phrasing, in-house abbreviations, and typos—not questions neatly rephrased by the person who wrote the documentation.

  1. 01
    Start with the real questions
    Thirty to fifty questions based on real usage are better than two hundred invented questions. Include the ones that failed: they are the most instructive.
  2. 02
    Write the reference answer
    For each question, the correct answer, written in one or two sentences. Ragas uses it to estimate context recall: the LLM-based version treats this reference as a substitute for the expected passages, avoiding the need to annotate passages one by one.
  3. 03
    Keep unanswered cases
    Questions the corpus doesn't answer. A good system should say so; without these cases, you can never measure that aspect of quality.
  4. 04
    Freeze the game
    It must not change at the same time as the system; otherwise, no comparison over time is possible.

Ragas also offers assisted generation of synthetic test sets from the corpus itself: the documentation distinguishes one-hop questions (a single source) from multi-hop questions (multiple sources to connect), both specific and abstract. An official guide shows how to adapt it to a non-English corpus; the example uses Spanish, providing a starting point for an initial evaluation before enough real user questions have accumulated. It’s a convenient starting point, never a substitute: a fully synthetic set misses the awkward phrasing and typos that specifically reveal the real weaknesses of a document search system.

#Run everything locally

Ragas relies on two models for evaluation: a judge model that reads the question, passages, and answer, and an embedding model for similarity measurements. Ragas's quickstart uses OpenAI by default, but shows the Ollama variant: an OpenAI-compatible client pointed at http://localhost:11434/v1, passed to the llm_factory function. Only answer relevance requires an embedding model; faithfulness, precision, and context recall use only the judge.

Two conditions make a local judge credible. The first is structured output: current metrics require the judge to make intermediate decisions in a prescribed format (claim extraction, verdicts). A local model may respond in prose and fail to meet this contract, producing JSON errors or empty scores (NaN); a OneUptime guide recommends testing the judge with a tiny call before any campaign. No official source sets a minimum model size: you must verify it yourself. The second is context: Ollama applies 4,000 tokens of context by default below 24 GiB of VRAM, and the documentation recommends increasing the value (OLLAMA_CONTEXT_LENGTH variable) for heavy workloads. A judge whose context overflows is effectively scoring truncated text.

Python: a local judge (current Ragas API)
# pip install ragas openai
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness, ContextRecall

# Client compatible OpenAI pointé sur Ollama (la clé est un simple libellé)
client = AsyncOpenAI(api_key="ollama", base_url="http://localhost:11434/v1")
juge = llm_factory("mistral", provider="openai", client=client)  # nom du modèle tiré avec ollama pull

fidelite = Faithfulness(llm=juge)
resultat = fidelite.score(
    user_input="Où se trouve la tour Eiffel ?",
    response="La tour Eiffel se trouve à Paris.",
    retrieved_contexts=["La tour Eiffel est située à Paris, en France."],
)
print(resultat.value)  # entre 0 et 1
→
Replay rather than restart at random
Keep each campaign in a dated file with the tested system version and the Ragas version. That lets you answer “when did context precision start to decline?” six months later instead of rediscovering it from a user complaint.
!
Legacy API, telemetry
Many tutorials still use LangchainLLMWrapper with the evaluate function: that’s the old API, which the documentation says is deprecated in 0.4 and removed in 1.0. Another point: Ragas collects anonymized usage data by default; you can disable it by setting the RAGAS_DO_NOT_TRACK variable to true. For an evaluation that must remain offline, enable it.

#Read the results accurately

These are indicators, not scores
A score of 0.82 does not mean “82% correct answers.” This figure is used to compare two versions of the same system, not to certify absolute quality.
The judge has biases
The reference paper on LLM judges (arXiv 2306.05685) describes position, verbosity, and self-preference biases. Use the same judge from one campaign to the next; otherwise, the differences measure the judge rather than the system.
A single campaign proves nothing
Generation is variable. On a small test set, a difference of a few hundredths may simply be noise.
Read a few cases manually
The numbers tell you where to look; they do not tell you what is wrong. The ten worst cases in a campaign teach you more than the overall average.
Two ways to judge it, and what they are worth
MethodWhat it offersIts limitation
Human review for a few casesThe only real quality control for a high-stakes subjectDoes not scale; requires expert time for every campaign
Local judge model (Ragas)Reproducible, free after installation, quickly compares two versionsPosition, verbosity, and self-preference bias; not an absolute truth score
User rating (thumbs up)Reflects real-world usage, free to collectOften little feedback, and a thumbs-up doesn’t indicate which step failed

#Concrete use cases

Choose a chunk size
Replaying the same test set with chunks of 256, 512, and then 1024 tokens gives a recall figure for each configuration, rather than an unverified preference.
Validate an embedding model change
A new embedding model, even if announced as better on a general benchmark, can reduce context accuracy on your specific corpus; only a local evaluation campaign can show this.
Justify adding a reranker
Comparing context precision before and after reranking puts a number on a gain that would otherwise remain a shared impression in a meeting.
Track a regression after updating the corpus
Adding new documents can dilute search results; replay the frozen test set after each import détecte to catch degradation before a user reports it.
Choosing between two providers or architectures
When faced with two competing proposals for building the same document pipeline, a score obtained on the same test set and corpus settles the matter faster than a sales demonstration.

#Organize campaigns

Campaign frequency
After every change to chunking, the embedding model, or the generation model. There is no need for a daily campaign on a stable system.
Test corpus size
The question set remains small; it is the evaluated document corpus that must be a realistic copy—or the entirety—of the production corpus.
The true cost of a campaign
Each metric calls the judge several times (for faithfulness: extracting the claims, then verifying each one). The cost therefore grows with the number of questions multiplied by the number of metrics: time five questions before launching the fifty, then plan the campaign accordingly.

#The method's limitations

Evaluating with a model means asking an artificial intelligence to judge another artificial intelligence: the method is useful, economical, and imperfect. It detects regressions and ranks variants; it does not replace human review in a high-stakes domain, where a factual error carries responsibility. Finally, each campaign consumes compute time: on a machine that also serves users, run it when the load is low—at night or over a weekend—rather than during peak usage hours.

The verbosity bias deserves a little more explanation because it is counterintuitive. A OneUptime article explains that, in a long response, a judge has more opportunities to quote the rubric's keywords, appear exhaustive, and offer a convincing justification. A prompt that pushes the evaluated model to be more verbose can therefore raise a score without improving the user's answer. For faithfulness, the Ragas documentation also mentions a variant based on HHEM-2.1-Open, a small free hallucination-detection classifier from Vectara: it replaces the verification step with a specialized model, worth trying if your local judge is unstable.

The simplest solution is methodological: never compare a faithfulness score obtained with one judge to a score obtained with another, or to a score published by a third party. The number only makes sense within the same measurement setup — same judge, same evaluation prompt, same metrics version. It is an internal comparison tool, not a universal ranking.

#FAQ

Can Ragas work without a cloud model?+
Yes, provided you explicitly configure a local judge. The library, licensed under Apache 2.0 and boasting more than 15,000 stars on GitHub, uses OpenAI by default in its quickstart, but its guide shows an OpenAI-compatible client pointed at Ollama. An embedding model is required only for answer relevance. Enable RAGAS_DO_NOT_TRACK to disable telemetry.
How many questions do you need in a test set?+
Thirty to fifty real questions already provide a usable basis for detecting a clear regression. Beyond that, value comes from the diversity of cases—including questions with no answer in the corpus—far more than from the raw number of questions asked of the tested document pipeline.
Which metric should you look at first?+
Faithfulness, because a fabricated answer costs more than no answer in a professional context where a mistake entails liability. Then context recall, which tells you whether the information was merely available when the answer was generated, before even judging how the answer was phrased.
Can you compare two models with Ragas?+
Yes, it is one of its most useful applications: the same set of questions, the same corpus, the same evaluator—you change only the generation model. It is the only reliable way to determine whether a larger model improves your answers, on your actual documents.
Is a local judge model reliable?+
Enough to compare two versions, provided it follows the structured output format requested by Ragas (test it in an isolated call) and its context contains everything it needs to read: Ollama applies 4,000 tokens by default under 24 GiB of VRAM. It does not replace human review for high-stakes topics, and it has position, verbosity, and self-preference biases.
Can Ragas automatically generate a test set?+
Yes, the library provides assisted generation of synthetic questions from the corpus, which is useful for starting an initial evaluation before enough real questions have accumulated. It does not replace questions actually asked by users, whose awkward phrasing reveals weaknesses that a clean, synthetic dataset never exposes in the same way.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.