RAG over case law Légifrance
For a RAG system over case law, download DILA's open XML archives (CASS for decisions published in the Bulletin of the Cour de cassation, INCA for unpublished decisions), split each decision by section, index them with BGE-M3 in Qdrant, and have it cite the appeal number and ECLI. The archived data weighs a few hundred MB compressed, not tens of GB.
French case law is published as open data, but its datasets have a specific scope, a distinctive XML structure, and pitfalls that generic tutorials overlook. This guide builds a local semantic search engine for these decisions: downloading, reading the XML, splitting by section, indexing, filtered search, and limitations to display to users.
#What a legal RAG is and what we ask of it
A legal RAG is a semantic search engine for court decisions, paired with a model that drafts an answer from the passages it retrieves. The search turns the question into a vector, finds the closest excerpts in a database of decisions, and then the model synthesizes them while citing their references. Compared with a model alone, this is decisive in law: the model does not answer from memory; it answers from texts you can review, with the jurisdiction, date, appeal number, and ECLI identifier.
This guide builds the engine with open data published by the Directorate of Legal and Administrative Information (DILA), on a local machine: no legal question is sent to an external service. The use case is a question such as “termination of a fixed-term contract: what recent decisions has the French Court of Cassation issued?”, with a list of sourced passages in response. The system does not provide an opinion: it retrieves and cites, and it is up to the professional to assess the result.
#Official datasets: what they really contain
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
DILA publishes decisions in separate databases, each with its own scope. An important point that is often misunderstood: these databases do not contain every decision issued in France. The Court of Cassation collection published under the name CASS includes decisions published in the Bulletin, those of the civil chambers since 1960, and those of the criminal chamber since 1963, with headings and summaries written by the judges. Unpublished decisions, not published in the Bulletin, are in a separate database, INCA, distributed since 1989.
| Base | Content | Stock archive (compressed) |
|---|---|---|
| CASS | Decisions of the French Court of Cassation published in the Bulletin (civil since 1960, criminal since 1963) | approximately 248 MB |
| INCA | Unpublished decisions of the Court of Cassation, not published in the Bulletin, since 1989 | about 655 MB |
| CAPP | Selection of civil and criminal decisions from appellate courts and courts of first instance | about 279 MB |
| JADE | Conseil d'État, administrative courts of appeal, Tribunal des conflits (selected according to jurisdiction) | about 1.2 GB |
The sizes above are those of the full archives listed on the DILA server when checked on September 30, 2026: a few hundred megabytes compressed per database, nowhere near the tens of gigabytes sometimes cited. The uncompressed volume is larger because each decision is a small XML file, but a standard workstation is sufficient. Measure it on your disk before planning.
There is a second route: the Judilibre API, implemented by the Cour de cassation to make a free, open database available to the public, populated with publicly issued decisions that may be enriched and pseudonymized. It uses an authenticated programming interface, whereas the DILA archives are simply files to download. For a local RAG system, the archives are the simplest starting point; Judilibre becomes relevant when you want a more complete and up-to-date corpus.
#Download the foundation: initial release, then updates
Each database follows the same pattern on the DILA server: a stock archive whose name starts with Freemium and ends with global, followed by incremental archives containing the new material. As of the consultation date, the Cour de cassation stock archive was dated July 13, 2025, and weekly archives supplemented it through the end of September 2026. You must therefore download the stock archive, then apply all subsequent incremental archives in order.
The name of the stock archive changes when DILA regenerates it: review the directory listing before hard-coding a name. Also avoid downloading the directory repeatedly: one stock archive and the incremental updates are enough, and a recursive mirror unnecessarily overloads the public server.
#Read DILA XML without selecting the wrong field
The format is a family of document-type definitions shared by several databases, published by DILA under the name DTD Légifrance. In a weekly archive from September 2026, the structure of a Court of Cassation ruling is as follows: a shared metadata block (identifier, type), a legal metadata block (title, decision date, court, ruling), and a block specific to the judicial system (case number, panel, ECLI, Bulletin publication indicator). The full text is in the text block, under the content element, with line breaks encoded as br tags.
Two pitfalls appear in tutorials. First, the NUMERO field in the legal block is an internal number; the appeal number—the one cited by legal professionals—is in the specific block under NUMEROS_AFFAIRES. Second, retrieving all the file’s text with itertext mixes metadata into the judgment body and contaminates the vectors: you need to target the content element.
#Break it down by decision structure, not token count
A recent decision by the French Court of Cassation follows a recognizable structure. In the rulings examined, you find the headings “Facts and procedure,” “Review of the grounds,” then, for each ground, “Statement of the ground” and “Court’s response,” and finally the operative part introduced by “ON THESE GROUNDS.” Splitting every 700 tokens separates the question posed to the Court from its answer and produces ambiguous excerpts. First split according to these headings, then split only blocks that are too long.
A second improvement that costs little and delivers a lot: prefix each excerpt with its header (the decision title and ECLI). An isolated excerpt such as “the appeals court shifted the burden of proof” tells you neither which court it came from nor its date. With the header, both the embedding model and the generator have the context. For older decisions with a different structure, fall back to paragraph-based chunking with overlap.
#Index with BGE-M3 and Qdrant
BGE-M3 is a multilingual embedding model that produces 1,024-dimensional vectors and accepts inputs of up to 8,192 tokens, according to its official model card. It handles legal French correctly in common use cases; still, measure it on your own questions. Qdrant works well for storing vectors with their metadata. Its Python client offers a local serverless mode, but its documentation targets development, prototyping, and testing; for several hundred thousand excerpts, run the server (for example, with Docker).
The common tutorial example contains a silent bug: using the decision index as a point identifier overwrites all excerpts from the same decision except the last one. Each excerpt needs a unique identifier. Another detail: to filter by date, store an integer in the YYYYMMDD format in the metadata, enabling range filtering.
#Query with filters and read the results
The search encodes the question with the same model and asks Qdrant for the closest excerpts, optionally restricted by a filter. Filters on training, solution, or date are a real advantage over a traditional full-text search: « chambre sociale, depuis 2023 » becomes two metadata conditions, not one more word in the query.
#Why semantic-only search isn't enough for legal work
A lawyer often looks for exact details: a docket number, an article number, or an established legal expression. Semantic similarity retrieves these poorly because two similar numbers have no semantic relationship. The authors of BGE-M3 also recommend the following RAG pipeline in the model card: hybrid search followed by reranking. So add BM25-style lexical search alongside the vectors, merge the two lists, then pass the best candidates to a reranking model.
For a legal corpus, this addition is the most important improvement after good chunking. The guides on hybrid search and reranking cover the implementation; this guide focuses only on what is specific to case law.
#Generation, updates, and citation control
For generation, provide the model with the five to eight best excerpts along with their references, and require it to answer only from them. A 9-billion-parameter model such as Qwen 3.5 9B (6.6 GB in the Ollama library) is sufficient for summarizing excerpts; Mistral Small 24B (14 GB) requires 16 GB of video memory. The prompt must require citing the docket number and ECLI for every assertion, and answering “no decision found” when the excerpts do not answer the question.
For updates, DILA publishes incremental archives, approximately weekly for CASS and INCA based on the lists reviewed. A scheduled process should download the missing ones, analyze them, and add the excerpts, using stable identifiers to prevent duplicates. Finally, verify programmatically that every reference cited in the answer appears in the supplied excerpts: it is the same control logic used in the guide to contract analysis.
#Limitations to show users
- Partial coverage
- CASS contains only rulings published in the Bulletin, INCA contains unpublished rulings, and CAPP contains a selection of court of appeals rulings: the absence of a decision from the database does not prove that it does not exist.
- Unpublished or too-recent decisions
- A very recent decision may not yet be included in the latest incremental archive. Cross-check with Légifrance for a sensitive case.
- Evolution of the law
- An old ruling may have been superseded by a reversal or reform. The system retrieves text; it does not measure the current authority of a solution.
- Pseudonymization
- Names are redacted. Never try to identify the parties.
- No legal advice
- The tool helps find and cite sources. Fact qualification, applying them to the case, and advice remain the professional's responsibility.
- Analyze a contract without data leakage
- Hybrid search: BM25 and vectors
- Add a reranker to your pipeline
- Document chunking strategies
- BGE-M3: embeddings for French
- Qdrant: the vector database for local RAG
- Source: CASS dataset on data.gouv.fr
- Source: CASS archives on the DILA server
- Source: Judilibre API on data.gouv.fr
- Source: BGE-M3 model sheet
- Source: Qdrant Python client
What is a legal RAG?+
Where can you find Court of Cassation case law in open data?+
Is the corpus complete?+
Which embedding model should you choose for French legal text?+
Do you need a GPU to index decisions?+
Can these decisions be used in a commercial product?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.