Ragas: evaluate your local RAG with chiffres
Ragas is an open-source Python library (Apache 2.0 license, more than 15,000 stars on GitHub) that quantifies the quality of a document-retrieval pipeline: faithfulness, answer relevance, context precision, and context recall. It can run entirely with a local judge model and embedding model, avoiding the need to send your documents and answers to a third-party service just to evaluate them.
You tune a document-retrieval pipeline blindly until you have a number: you change chunk sizes, replace the embedding model, and judge by feel on three questions. Ragas is a Python library that puts metrics behind these intuitions—and separates, when an answer is poor, retrieval’s fault from the model’s. It can run entirely with local models, avoiding the need to send your documents to a third-party service for evaluation.
#Why “it looks good” isn’t enough
Ragas is an open-source Python library under the Apache 2.0 license that quantifies the quality of a document-retrieval pipeline: the answer's faithfulness to the provided passages, its relevance to the question, and the precision and recall of the retrieved context. This separates retrieval errors from model errors. It can run with a local judge served by Ollama, provided that model follows the structured-output format Ragas requires and its context contains the entire question, passages, and answer. Before adopting it, keep three points in mind: its scores are for comparing two versions of the same system, not for rating absolute quality; the test set matters more than the library; and online tutorials mix the old and new APIs.
When an answer is wrong, two very different causes can hide behind the same symptom: either the retrieved passages did not contain the information, or they did and the model answered beside the point. The fix is not the same—chunking and embeddings in the first case, the model and prompt in the second. Without measurement, you are fixing things at random, and an improvement on three questions can worsen five others without you noticing.
The other pitfall is comparison. “Is the 27-billion-parameter model really better here?” is answered by rerunning the same set of questions, not by discussing it. That is what reproducible evaluation makes possible. Ragas has surpassed 15,000 stars on GitHub. The latest version published on PyPI is 0.4.3, from January 13, 2026, and the repository has changed GitHub organizations (vibrantlabsai). Pin the version you use.
#The four metrics that matter
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
| Metric | Question asked | What it blames when it drops | What you need to provide |
|---|---|---|---|
| Faithfulness | Is the response fully supported by the supplied passages? Score: supported claims divided by total claims in the response | The model makes things up or extrapolates | Question, answer, retrieved passages |
| Response relevance | Does the answer address the question asked? The judge generates three questions from the answer, then compares their similarity with the original question | The prompt or model goes off topic | Question, answer, and an embedding model |
| Context accuracy | Are the relevant passages ranked near the top? Average precision at each rank | Search returns noise or ranks results poorly | Question, passages in retrieved order, reference answer |
| Context reminder | Did we retrieve everything we needed? Share of the reference answer's claims supported by the passages | Chunks too small, threshold too strict, bad embeddings | Question, passages, reference answer |
Ragas documentation specifies that answer relevance does not judge correctness: it only measures how well the answer fits the question, and penalizes incomplete answers or ones full of unnecessary details. That is why it does not replace faithfulness. Cross-reading is what makes these figures useful. Low recall with high faithfulness describes an honest but poorly supplied system: you need to improve retrieval. Low faithfulness with good recall describes the opposite: the information was there, but the model embellished. The two are corrected at opposite points in the pipeline.
#Build a test set
This is the part that requires work, and it can’t be automated without consequences. A useful dataset contains questions users actually asked, including their awkward phrasing, in-house abbreviations, and typos—not questions neatly rephrased by the person who wrote the documentation.
- 01Start with the real questionsThirty to fifty questions based on real usage are better than two hundred invented questions. Include the ones that failed: they are the most instructive.
- 02Write the reference answerFor each question, the correct answer, written in one or two sentences. Ragas uses it to estimate context recall: the LLM-based version treats this reference as a substitute for the expected passages, avoiding the need to annotate passages one by one.
- 03Keep unanswered casesQuestions the corpus doesn't answer. A good system should say so; without these cases, you can never measure that aspect of quality.
- 04Freeze the gameIt must not change at the same time as the system; otherwise, no comparison over time is possible.
Ragas also offers assisted generation of synthetic test sets from the corpus itself: the documentation distinguishes one-hop questions (a single source) from multi-hop questions (multiple sources to connect), both specific and abstract. An official guide shows how to adapt it to a non-English corpus; the example uses Spanish, providing a starting point for an initial evaluation before enough real user questions have accumulated. It’s a convenient starting point, never a substitute: a fully synthetic set misses the awkward phrasing and typos that specifically reveal the real weaknesses of a document search system.
#Run everything locally
Ragas relies on two models for evaluation: a judge model that reads the question, passages, and answer, and an embedding model for similarity measurements. Ragas's quickstart uses OpenAI by default, but shows the Ollama variant: an OpenAI-compatible client pointed at http://localhost:11434/v1, passed to the llm_factory function. Only answer relevance requires an embedding model; faithfulness, precision, and context recall use only the judge.
Two conditions make a local judge credible. The first is structured output: current metrics require the judge to make intermediate decisions in a prescribed format (claim extraction, verdicts). A local model may respond in prose and fail to meet this contract, producing JSON errors or empty scores (NaN); a OneUptime guide recommends testing the judge with a tiny call before any campaign. No official source sets a minimum model size: you must verify it yourself. The second is context: Ollama applies 4,000 tokens of context by default below 24 GiB of VRAM, and the documentation recommends increasing the value (OLLAMA_CONTEXT_LENGTH variable) for heavy workloads. A judge whose context overflows is effectively scoring truncated text.
- Build the RAG pipeline you are going to evaluate
- Choosing an embedding model for French
- Add a reranker when context accuracy is low
- Langfuse: track and store the scores obtained
#Read the results accurately
- These are indicators, not scores
- A score of 0.82 does not mean “82% correct answers.” This figure is used to compare two versions of the same system, not to certify absolute quality.
- The judge has biases
- The reference paper on LLM judges (arXiv 2306.05685) describes position, verbosity, and self-preference biases. Use the same judge from one campaign to the next; otherwise, the differences measure the judge rather than the system.
- A single campaign proves nothing
- Generation is variable. On a small test set, a difference of a few hundredths may simply be noise.
- Read a few cases manually
- The numbers tell you where to look; they do not tell you what is wrong. The ten worst cases in a campaign teach you more than the overall average.
| Method | What it offers | Its limitation |
|---|---|---|
| Human review for a few cases | The only real quality control for a high-stakes subject | Does not scale; requires expert time for every campaign |
| Local judge model (Ragas) | Reproducible, free after installation, quickly compares two versions | Position, verbosity, and self-preference bias; not an absolute truth score |
| User rating (thumbs up) | Reflects real-world usage, free to collect | Often little feedback, and a thumbs-up doesn’t indicate which step failed |
#Concrete use cases
- Choose a chunk size
- Replaying the same test set with chunks of 256, 512, and then 1024 tokens gives a recall figure for each configuration, rather than an unverified preference.
- Validate an embedding model change
- A new embedding model, even if announced as better on a general benchmark, can reduce context accuracy on your specific corpus; only a local evaluation campaign can show this.
- Justify adding a reranker
- Comparing context precision before and after reranking puts a number on a gain that would otherwise remain a shared impression in a meeting.
- Track a regression after updating the corpus
- Adding new documents can dilute search results; replay the frozen test set after each import détecte to catch degradation before a user reports it.
- Choosing between two providers or architectures
- When faced with two competing proposals for building the same document pipeline, a score obtained on the same test set and corpus settles the matter faster than a sales demonstration.
#Organize campaigns
- Campaign frequency
- After every change to chunking, the embedding model, or the generation model. There is no need for a daily campaign on a stable system.
- Test corpus size
- The question set remains small; it is the evaluated document corpus that must be a realistic copy—or the entirety—of the production corpus.
- The true cost of a campaign
- Each metric calls the judge several times (for faithfulness: extracting the claims, then verifying each one). The cost therefore grows with the number of questions multiplied by the number of metrics: time five questions before launching the fifty, then plan the campaign accordingly.
#The method's limitations
Evaluating with a model means asking an artificial intelligence to judge another artificial intelligence: the method is useful, economical, and imperfect. It detects regressions and ranks variants; it does not replace human review in a high-stakes domain, where a factual error carries responsibility. Finally, each campaign consumes compute time: on a machine that also serves users, run it when the load is low—at night or over a weekend—rather than during peak usage hours.
The verbosity bias deserves a little more explanation because it is counterintuitive. A OneUptime article explains that, in a long response, a judge has more opportunities to quote the rubric's keywords, appear exhaustive, and offer a convincing justification. A prompt that pushes the evaluated model to be more verbose can therefore raise a score without improving the user's answer. For faithfulness, the Ragas documentation also mentions a variant based on HHEM-2.1-Open, a small free hallucination-detection classifier from Vectara: it replaces the verification step with a specialized model, worth trying if your local judge is unstable.
The simplest solution is methodological: never compare a faithfulness score obtained with one judge to a score obtained with another, or to a score published by a third party. The number only makes sense within the same measurement setup — same judge, same evaluation prompt, same metrics version. It is an internal comparison tool, not a universal ranking.
- Source: official Ragas repository on GitHub
- Source: verbosity bias in automated judges
- Source: official documentation for the faithfulness metric
- Source: Ragas quickstart, Ollama variant
- Source: default context length of Ollama
- Source: evaluating a RAG with a local judge that produces invalid JSON
- Source: reference article on LLM judges
#FAQ
Can Ragas work without a cloud model?+
How many questions do you need in a test set?+
Which metric should you look at first?+
Can you compare two models with Ragas?+
Is a local judge model reliable?+
Can Ragas automatically generate a test set?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.