Advanced 12 minLegal

RAG over case law Légifrance

Direct response

For a RAG system over case law, download DILA's open XML archives (CASS for decisions published in the Bulletin of the Cour de cassation, INCA for unpublished decisions), split each decision by section, index them with BGE-M3 in Qdrant, and have it cite the appeal number and ECLI. The archived data weighs a few hundred MB compressed, not tens of GB.

French case law is published as open data, but its datasets have a specific scope, a distinctive XML structure, and pitfalls that generic tutorials overlook. This guide builds a local semantic search engine for these decisions: downloading, reading the XML, splitting by section, indexing, filtered search, and limitations to display to users.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What a legal RAG is and what we ask of it

A legal RAG is a semantic search engine for court decisions, paired with a model that drafts an answer from the passages it retrieves. The search turns the question into a vector, finds the closest excerpts in a database of decisions, and then the model synthesizes them while citing their references. Compared with a model alone, this is decisive in law: the model does not answer from memory; it answers from texts you can review, with the jurisdiction, date, appeal number, and ECLI identifier.

This guide builds the engine with open data published by the Directorate of Legal and Administrative Information (DILA), on a local machine: no legal question is sent to an external service. The use case is a question such as “termination of a fixed-term contract: what recent decisions has the French Court of Cassation issued?”, with a list of sourced passages in response. The system does not provide an opinion: it retrieves and cites, and it is up to the professional to assess the result.

i
What sets this guide apart
Most RAG tutorials start with documents you own. Here, the challenge lies elsewhere: understanding what public datasets actually contain, their XML structure, volume, update frequency, and coverage limitations. Every figure in the next section was taken from official sources.

#Official datasets: what they really contain

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

DILA publishes decisions in separate databases, each with its own scope. An important point that is often misunderstood: these databases do not contain every decision issued in France. The Court of Cassation collection published under the name CASS includes decisions published in the Bulletin, those of the civil chambers since 1960, and those of the criminal chamber since 1963, with headings and summaries written by the judges. Unpublished decisions, not published in the Bulletin, are in a separate database, INCA, distributed since 1989.

DILA case-law databases (according to the data.gouv.fr data sheets)
BaseContentStock archive (compressed)
CASSDecisions of the French Court of Cassation published in the Bulletin (civil since 1960, criminal since 1963)approximately 248 MB
INCAUnpublished decisions of the Court of Cassation, not published in the Bulletin, since 1989about 655 MB
CAPPSelection of civil and criminal decisions from appellate courts and courts of first instanceabout 279 MB
JADEConseil d'État, administrative courts of appeal, Tribunal des conflits (selected according to jurisdiction)about 1.2 GB

The sizes above are those of the full archives listed on the DILA server when checked on September 30, 2026: a few hundred megabytes compressed per database, nowhere near the tens of gigabytes sometimes cited. The uncompressed volume is larger because each decision is a small XML file, but a standard workstation is sufficient. Measure it on your disk before planning.

There is a second route: the Judilibre API, implemented by the Cour de cassation to make a free, open database available to the public, populated with publicly issued decisions that may be enriched and pseudonymized. It uses an authenticated programming interface, whereas the DILA archives are simply files to download. For a local RAG system, the archives are the simplest starting point; Judilibre becomes relevant when you want a more complete and up-to-date corpus.

!
Personal data: you are responsible
DILA specifies on its fact sheets that making datasets potentially containing personal data available does not exempt the reuser from complying with the Informatique et Libertés law. Decisions are pseudonymized (passages such as [V] [S] replace names), but never attempt to re-identify individuals, and check the scope of your project with your data protection officer.

#Download the foundation: initial release, then updates

Each database follows the same pattern on the DILA server: a stock archive whose name starts with Freemium and ends with global, followed by incremental archives containing the new material. As of the consultation date, the Cour de cassation stock archive was dated July 13, 2025, and weekly archives supplemented it through the end of September 2026. You must therefore download the stock archive, then apply all subsequent incremental archives in order.

Download the CASS stock, then the updates
mkdir -p dila/CASS && cd dila/CASS
curl -O https://echanges.dila.gouv.fr/OPENDATA/CASS/Freemium_cass_global_20250713-140000.tar.gz
tar xzf Freemium_cass_global_20250713-140000.tar.gz

# Puis chaque archive incrémentale, dans l'ordre chronologique
curl -s https://echanges.dila.gouv.fr/OPENDATA/CASS/ | grep -o 'CASS_2026[0-9-]*\.tar\.gz' | sort -u > maj.txt
for f in $(cat maj.txt); do curl -sO https://echanges.dila.gouv.fr/OPENDATA/CASS/$f && tar xzf $f; done

The name of the stock archive changes when DILA regenerates it: review the directory listing before hard-coding a name. Also avoid downloading the directory repeatedly: one stock archive and the incremental updates are enough, and a recursive mirror unnecessarily overloads the public server.

#Read DILA XML without selecting the wrong field

The format is a family of document-type definitions shared by several databases, published by DILA under the name DTD Légifrance. In a weekly archive from September 2026, the structure of a Court of Cassation ruling is as follows: a shared metadata block (identifier, type), a legal metadata block (title, decision date, court, ruling), and a block specific to the judicial system (case number, panel, ECLI, Bulletin publication indicator). The full text is in the text block, under the content element, with line breaks encoded as br tags.

Two pitfalls appear in tutorials. First, the NUMERO field in the legal block is an internal number; the appeal number—the one cited by legal professionals—is in the specific block under NUMEROS_AFFAIRES. Second, retrieving all the file’s text with itertext mixes metadata into the judgment body and contaminates the vectors: you need to target the content element.

Parse a CASS decision (structure verified against a 2026 archive)
from lxml import etree
from pathlib import Path
import re

def parser(chemin):
    racine = etree.parse(str(chemin)).getroot()

    def val(xp):
        e = racine.find(xp)
        return (e.text or '').strip() if e is not None else None

    contenu = racine.find('.//TEXTE/BLOC_TEXTUEL/CONTENU')
    if contenu is None:
        return None
    for br in contenu.iter('br'):
        br.tail = '\n' + (br.tail or '')
    texte = ''.join(contenu.itertext())
    texte = re.sub(r'[ \t]+', ' ', re.sub(r'\n\s*\n+', '\n', texte)).strip()
    return {
        'id': val('.//META_COMMUN/ID'),
        'titre': val('.//META_JURI/TITRE'),
        'date': val('.//META_JURI/DATE_DEC'),
        'juridiction': val('.//META_JURI/JURIDICTION'),
        'solution': val('.//META_JURI/SOLUTION'),
        'pourvoi': val('.//META_JURI_JUDI/NUMEROS_AFFAIRES/NUMERO_AFFAIRE'),
        'formation': val('.//META_JURI_JUDI/FORMATION'),
        'ecli': val('.//META_JURI_JUDI/ECLI'),
        'texte': texte,
    }

def decisions(dossier):
    for chemin in Path(dossier).rglob('*.xml'):
        try:
            d = parser(chemin)
            if d:
                yield d
        except etree.XMLSyntaxError as e:
            print('ignoré', chemin, e)
→
Adapt to each base model
These paths were identified in CASS. The JADE, CAPP, and INCA databases share the Légifrance DTD but have different database-specific blocks: open two or three files from each database before writing the parser, and test it on a sample of one hundred decisions.

#Break it down by decision structure, not token count

A recent decision by the French Court of Cassation follows a recognizable structure. In the rulings examined, you find the headings “Facts and procedure,” “Review of the grounds,” then, for each ground, “Statement of the ground” and “Court’s response,” and finally the operative part introduced by “ON THESE GROUNDS.” Splitting every 700 tokens separates the question posed to the Court from its answer and produces ambiguous excerpts. First split according to these headings, then split only blocks that are too long.

A second improvement that costs little and delivers a lot: prefix each excerpt with its header (the decision title and ECLI). An isolated excerpt such as “the appeals court shifted the burden of proof” tells you neither which court it came from nor its date. With the header, both the embedding model and the generator have the context. For older decisions with a different structure, fall back to paragraph-based chunking with overlap.

Split by section with heading
import re

COUPURE = re.compile(r"\n(?=(?:Faits et procédure|Examen des moyens|Sur le |Énoncé du moyen|Enoncé du moyen|Réponse de la Cour|PAR CES MOTIFS))")

def extraits(d, max_car=2200):
    en_tete = f"{d['titre']} ({d['ecli']})\n"
    sortie, tampon = [], ''
    for bloc in COUPURE.split(d['texte']):
        for para in bloc.split('\n'):
            if len(tampon) + len(para) > max_car and tampon:
                sortie.append(en_tete + tampon)
                tampon = ''
            tampon += para + '\n'
    if tampon.strip():
        sortie.append(en_tete + tampon)
    return sortie

#Index with BGE-M3 and Qdrant

BGE-M3 is a multilingual embedding model that produces 1,024-dimensional vectors and accepts inputs of up to 8,192 tokens, according to its official model card. It handles legal French correctly in common use cases; still, measure it on your own questions. Qdrant works well for storing vectors with their metadata. Its Python client offers a local serverless mode, but its documentation targets development, prototyping, and testing; for several hundred thousand excerpts, run the server (for example, with Docker).

The common tutorial example contains a silent bug: using the decision index as a point identifier overwrites all excerpts from the same decision except the last one. Each excerpt needs a unique identifier. Another detail: to filter by date, store an integer in the YYYYMMDD format in the metadata, enabling range filtering.

Batch indexing
from sentence_transformers import SentenceTransformer
from qdrant_client import QdrantClient, models

emb = SentenceTransformer('BAAI/bge-m3', device='cuda')
emb.max_seq_length = 1024
client = QdrantClient(host='localhost', port=6333)
if not client.collection_exists('cass'):
    client.create_collection('cass', vectors_config=models.VectorParams(
        size=1024, distance=models.Distance.COSINE))

def indexer(dossier, taille_lot=64):
    n, lot = 0, []
    def vider():
        vecs = emb.encode([x[0] for x in lot], batch_size=32, normalize_embeddings=True)
        client.upsert('cass', points=[models.PointStruct(id=x[2], vector=v.tolist(), payload=x[1])
                                       for x, v in zip(lot, vecs)])
        lot.clear()
    for d in decisions(dossier):
        for texte in extraits(d):
            n += 1
            payload = {k: d[k] for k in ('titre', 'ecli', 'pourvoi', 'juridiction', 'formation', 'solution')}
            payload['date_int'] = int(d['date'].replace('-', ''))
            payload['texte'] = texte
            lot.append((texte, payload, n))
            if len(lot) >= taille_lot:
                vider()
    if lot:
        vider()

The search encodes the question with the same model and asks Qdrant for the closest excerpts, optionally restricted by a filter. Filters on training, solution, or date are a real advantage over a traditional full-text search: « chambre sociale, depuis 2023 » becomes two metadata conditions, not one more word in the query.

Filtered search
def chercher(question, depuis=None, formation=None, k=8):
    conds = []
    if depuis:
        conds.append(models.FieldCondition(key='date_int', range=models.Range(gte=depuis)))
    if formation:
        conds.append(models.FieldCondition(key='formation', match=models.MatchValue(value=formation)))
    vec = emb.encode(question, normalize_embeddings=True).tolist()
    res = client.query_points('cass', query=vec, limit=k,
                              query_filter=models.Filter(must=conds) if conds else None)
    return res.points

for p in chercher('rupture anticipée du CDD par l employeur', depuis=20230101):
    print(p.payload['titre'], p.payload['pourvoi'], round(p.score, 3))

#Why semantic-only search isn't enough for legal work

A lawyer often looks for exact details: a docket number, an article number, or an established legal expression. Semantic similarity retrieves these poorly because two similar numbers have no semantic relationship. The authors of BGE-M3 also recommend the following RAG pipeline in the model card: hybrid search followed by reranking. So add BM25-style lexical search alongside the vectors, merge the two lists, then pass the best candidates to a reranking model.

For a legal corpus, this addition is the most important improvement after good chunking. The guides on hybrid search and reranking cover the implementation; this guide focuses only on what is specific to case law.

#Generation, updates, and citation control

For generation, provide the model with the five to eight best excerpts along with their references, and require it to answer only from them. A 9-billion-parameter model such as Qwen 3.5 9B (6.6 GB in the Ollama library) is sufficient for summarizing excerpts; Mistral Small 24B (14 GB) requires 16 GB of video memory. The prompt must require citing the docket number and ECLI for every assertion, and answering “no decision found” when the excerpts do not answer the question.

Synthesis system prompt
Tu es un assistant de recherche en jurisprudence. Tu réponds uniquement à partir des
extraits fournis. Pour chaque affirmation, cite entre crochets le numéro de pourvoi
et l'ECLI de la décision. Si les extraits ne permettent pas de répondre, écris :
aucune décision retrouvée. Tu ne donnes pas d'avis juridique.

For updates, DILA publishes incremental archives, approximately weekly for CASS and INCA based on the lists reviewed. A scheduled process should download the missing ones, analyze them, and add the excerpts, using stable identifiers to prevent duplicates. Finally, verify programmatically that every reference cited in the answer appears in the supplied excerpts: it is the same control logic used in the guide to contract analysis.

#Limitations to show users

Partial coverage
CASS contains only rulings published in the Bulletin, INCA contains unpublished rulings, and CAPP contains a selection of court of appeals rulings: the absence of a decision from the database does not prove that it does not exist.
Unpublished or too-recent decisions
A very recent decision may not yet be included in the latest incremental archive. Cross-check with Légifrance for a sensitive case.
Evolution of the law
An old ruling may have been superseded by a reversal or reform. The system retrieves text; it does not measure the current authority of a solution.
Pseudonymization
Names are redacted. Never try to identify the parties.
No legal advice
The tool helps find and cite sources. Fact qualification, applying them to the case, and advice remain the professional's responsibility.
FAQ
What is a legal RAG?+
It is a system that finds excerpts close to a question in a database of decisions or texts, then has a model draft an answer based solely on those excerpts, along with their references. Its value in law lies in traceability: every claim points to a decision that the professional can reread.
Where can you find Court of Cassation case law in open data?+
On the DILA server, in the CASS archives for rulings published in the Bulletin and INCA for unpublished rulings, or through the Cour de cassation's Judilibre API. The records are available on data.gouv.fr. The archives are simple XML files that are easy to process locally.
Is the corpus complete?+
No. CASS contains rulings published in the Bulletin, INCA contains unpublished rulings, CAPP contains a selection of appellate court decisions, and JADE contains a selection for administrative justice. For a broader corpus, the Cour de cassation offers Judilibre. Never claim that a decision does not exist just because it is missing from your database.
Which embedding model should you choose for French legal text?+
BGE-M3 is a common choice: multilingual, with 1,024-dimensional vectors and inputs of up to 8,192 tokens according to its specifications. No general ranking replaces testing: create twenty questions for legal professionals with the expected decisions, then compare the recall of several models on this sample.
Do you need a GPU to index decisions?+
It isn’t required, but encoding hundreds of thousands of excerpts is significantly faster with a graphics card. Without a GPU, indexing still works: schedule it overnight, in batches, and do it only once. Once the database is built, searching for an individual question remains fast on the CPU, and weekly updates are lightweight.
Can these decisions be used in a commercial product?+
The data.gouv.fr data sheets specify the license for each dataset, generally an open license that permits reuse with attribution. DILA nevertheless reminds us that the reuser remains subject to the French Data Protection and Civil Liberties Act. Check the license and have your data protection officer approve the project.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.