Beginner 11 minFrench

CroissantLLM and Lucie: the 100% French models in local

CroissantLLM, Lucie, Pleias, Mistral: the French open-weight LLM scene is very real, and several of these models run comfortably locally on a mainstream computer. This guide surveys the 2026 landscape of 100% French, self-hostable LLMs, compares their actual French-language quality with the best international models (Qwen 3.8, Gemma 4), and explains when the sovereignty argument justifies choosing Lucie over a Chinese or American model.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why a 100% French model?

Most high-performing open-weight LLMs in 2026 come from China (Qwen, DeepSeek, GLM) or the United States (Llama, Gemma, Phi). They often speak French very well — Qwen 3.8 and Gemma 4 are excellent, incidentally — but they were trained primarily on English and Chinese content, with cultural biases that show through: implicit references, awkwardly translated American phrasing, political perspectives, and an imperfect handling of formal and informal language.

A model trained in France addresses three distinct needs: native quality in formal French, culturally relevant content (law, administration, literature, media), and legal sovereignty in the sense that the weights, datasets, and developers fall under European law. For a law firm, government agency, or public research project, this is often decisive.

i
Open weights does not mean open source
On this topic, Pleias, CroissantLLM and Lucie publish the weights, training datasets, and code. That is rarely the case with Mistral, which remains open-weights but keeps its datasets closed. An important distinction if you want to audit or rederive a model.

#Overview of French LLMs in 2026

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Five families dominate today’s self-hostable French landscape. They are not in the same league in terms of size, quality, or licensing.

Mistral (Mistral AI, Paris)
The champion. Mistral Small 24B, Magistral, Devstral. Apache 2.0 for the open models, 128k context, French quality on par with the best cloud models.
CroissantLLM (Centrale Supélec / Illuin)
A truly open-source 1.3B bilingual FR-EN model: weights, datasets, and code. Designed for research and edge deployment.
Lucie (LINAGORA / OpenLLM-France)
7B Apache 2.0, community project, trained on French public supercomputers (Jean Zay). The first true French model of this size.
Pleias (Paris)
Family of 1B to 3B models trained exclusively on open data (Common Corpus). Strict compliance, perfect for the public sector.
Croco / CroissantCool / derivatives
Community fine-tunes for French instructions. Often based on Mistral or CroissantLLM, hosted on Hugging Face.

Alongside them, some non-French models stand out in French: Qwen 3.8 (Alibaba), Gemma 4 (Google), and gpt-oss (OpenAI). We put them into perspective below in the benchmarks.

#CroissantLLM: 1.3B bilingual, fully open

CroissantLLM was created through a collaboration between CentraleSupélec, Illuin Technology, Inria, and INRIA Paris. It is a 1.3-billion-parameter model trained on approximately 3 trillion tokens, split evenly between French and English. Small, certainly, but its distinctive feature is complete transparency: weights, datasets, training recipes, evaluation scripts—everything is published.

Size
1.3B parameters, ~800 MB in Q4. Runs on a CPU, a recent smartphone, or a Raspberry Pi 5.
Languages
Balanced FR-EN bilingual (1:1 in training).
License
MIT, commercial use permitted without restriction.
Variants
CroissantLLMBase (pretrained), CroissantLLMChat (instruction-tuned), CroissantCool (community RLHF).
Usage
Embedded, edge, research, local fine-tuning. Not a general-purpose assistant to pit against GPT-5.
Test CroissantLLM via Ollama (community GGUF variant)
ollama run croissantllm/CroissantLLMChat-v0.1-GGUF:Q4_K_M

>>> Explique-moi la différence entre une SARL et une SAS en 4 lignes.
→
Small model, big edge use
For truly local embedded AI — a home-automation assistant on Raspberry Pi, offline translation on a low-end laptop, or CPU-based ticket classification — CroissantLLM is one of the few French models at this scale. And it is good at it.

#Lucie: the French open-source bet

Lucie is LINAGORA's flagship project, developed in partnership with OpenLLM-France and several public research labs. Introduced in early 2025, it is a 7-billion-parameter model trained on the Jean Zay supercomputer at CNRS-IDRIS, on a corpus combining French, English, German, Spanish, Italian, and code, with strong weighting toward French.

Size
7B parameters, ~4.5 GB in Q4_K_M, ~5.5 GB in Q5_K_M.
VRAM
5 GB minimum in Q4. Runs comfortably on RTX 3060 12 GB, or on a 16 GB unified-memory Mac with an M-series chip.
License
Apache 2.0, so it can be used commercially and redistributed.
Context
32k tokens in chat version (Lucie-7B-Instruct).
Sovereignty
Weights hosted on HuggingFace, but the code and data are accessible. Self-hosted mirroring is possible for government agencies.

What stands out in use: Lucie handles French very cleanly, without the “English calque” still noticeable in Llama 4 with literary phrasing. It is, however, a notch below the best models in its tier, such as Qwen 3.5 9B, for code and mathematical reasoning.

Install Lucie via Ollama (community GGUF)
ollama pull OpenLLM-France/Lucie-7B-Instruct-gguf:Q4_K_M
ollama run OpenLLM-France/Lucie-7B-Instruct-gguf:Q4_K_M
i
Distribution Ollama
Depending on the version, Lucie may require a custom Modelfile to point to the GGUF downloaded from Hugging Face. The community maintains up-to-date recipes in the OpenLLM-France repository.

#Pleias: radical compliance through data

Pleias takes a rare approach: it trains its models only on Common Corpus—a public, traceable dataset with no copyrighted content. The result is 1B and 3B models that an administration can adopt without concern about intellectual-property issues, but whose capabilities remain modest compared with large models trained on the entire web.

Sizes
Pleias-1B and Pleias-3B (Pico, Nano versions). RAG-friendly.
Specialty
Good at RAG, classification, and structured extraction tasks. Weak at free-form chat.
Argument
No training on copyright-protected data. Supply chain auditing possible.
Typical use case
Summarizing public documents, processing administrative requests, and classifying mail.

Pleias is not a conversational assistant; it is an extraction and structuring tool. For a public-sector organization subject to the AI Act and a strict accountability chain, it is probably the only French model that meets all the requirements without compromise.

#Mistral: the French reference

No need to go longer: Mistral remains the French open-weights leader. Mistral Small 24B (Apache 2.0, ollama run mistral-small, 14 GB in Q4) is probably the best self-hostable French model in 2026 on enthusiast hardware. For code, Devstral 24B (Apache 2.0) is the in-house agent specialist. Mistral Magistral brings reasoning (chain of thought) in French, covered in a dedicated guide on the site.

Mistral Small 24B
~14 GB in Q4 (ollama run mistral-small). RTX 4090 or MacBook Pro M3/M4 with 32+ GB. Probably the best French/quality compromise.
Devstral 24B
~14 GB in Q4, Apache 2.0 (ollama run devstral:24b). The code-agent specialist from Mistral AI.
Mistral Magistral 24B
Chain-of-thought reasoning in French.
Codestral 22B
Code-specialized but non-production license: prohibited for work. Prefer Devstral 24B.
!
Watch out for recent licenses
Not everything published by Mistral is Apache 2.0. Mistral Large, recent Codestral, and certain Pixtral versions are under the Mistral Research License, which prohibits commercial use. Always check the license on the Hugging Face page before building a product on it.

#FR quality tests vs. Qwen 3.8 and Gemma 4

Here is a qualitative comparison grid based on typical uses by a French-speaking user (formal writing, PDF summarization, general legal questions, coding). Observations were conducted at temperature 0.3, 8k context, Q4_K_M, with identical prompts.

Mistral Small 24B (FR)
Excellent writing, polished register, rare agreement errors. Very good code. French reference.
gpt-oss 20B (US)
Excellent multilingual performance, with French exceeding expectations. Some English layers in literary text. Very fast (MXFP4).
Qwen 3.8 27B (CN)
Very good formal French, sometimes awkward in casual language. Code at Mistral's level, superior reasoning—consider setting reasoning to low, or it will overthink.
Lucie 7B (FR)
Clean French, with a fluent formal administrative register. Average coding ability. Good for writing and summarization.
Qwen 3.5 9B (CN)
Strong in French for its size (8 GB), 256k context, vision. A good writing and coding compromise on a small configuration.
Gemma 4 12B (US)
Good French, sometimes overly academic. Multimodal, good RAG, Apache 2.0.
CroissantLLM 1.3B (FR)
The level of a lightweight assistant; not in the same category. Excellent for its size.
→
Honest verdict
In 2026, the best self-hostable French model remains Mistral Small 24B. If legal sovereignty matters as much as quality, Lucie 7B is the best size / French / Apache 2.0 license compromise. If embedded deployment is the priority, it is CroissantLLM.

#Sovereignty and the AI Act

The European AI Act, which came into force in 2024–2026, imposes several obligations on providers and deployers of general-purpose AI models: technical documentation, a summary of training data, labeling of generated content, and systemic-risk assessments for very large models. Models trained and published in the EU simplify auditing, but the AI Act does not say that you need a French model—it says you must be able to document and explain the system.

Local hosting argument
Regardless of the model, running it locally resolves 90% of GDPR concerns. This matters more than the model's nationality.
Transparency argument
For an AI Act audit, a model whose datasets are public (Pleias, CroissantLLM, Lucie) is easier to defend than a model whose training is opaque.
Ecosystem argument
Supporting Mistral, LINAGORA, and Pleias helps consolidate a European industry. It is a political argument, but a real one for public procurement.
Legal argument
The model weights published under Apache 2.0 by a European entity reduce uncertainty around data transfers to the US.
i
Local solves the essentials
Running Qwen 3.8 27B on your own RTX 4090 on your premises, with no outbound connection, is very likely more GDPR-compliant than calling a Mistral API hosted in a public cloud, even a French one. The question “where does the model come from” comes after “where do the inferences run.”

#Which one to choose for your use case

  1. 01
    You're just getting started and want to test a French model
    Install Lucie 7B (ollama pull OpenLLM-France/Lucie-7B-Instruct-gguf:Q4_K_M), ~5 GB, a truly 100% French model that runs everywhere. For a stronger general-purpose French model, run ollama run mistral-small (Mistral Small 24B, 14 GB).
  2. 02
    You have 12 GB of VRAM (RTX 3060/4070)
    Lucie-7B-Instruct in Q5_K_M for pure French, or Qwen 3.5 9B in Q8 (11 GB) for a more capable generalist. You can also try Mistral Small 24B in Q3.
  3. 03
    You have 24 GB (RTX 4090/3090, Mac 32+ GB)
    Mistral Small 24B in Q4_K_M is probably the best self-hostable French LLM, period. Add Lucie 7B for fast “pure French” use cases.
  4. 04
    You are a government agency or public-sector organization
    Pleias 1B/3B for document RAG (zero copyright risk), Lucie 7B for general assistance, Mistral Small 24B when you need quality. Avoid models with unclear licensing.
  5. 05
    You use embedded / edge / Raspberry Pi
    CroissantLLM 1.3B Q4. It is the smallest French model that remains usable. See the LLM guide for Raspberry Pi for details.

#Go further

Three related paths depending on your next step: to go deeper on French reasoning, the guide dedicated to Mistral Magistral details the FR chain of thought; to build a RAG system over your PDFs with one of these models, the local RAG guide with Ollama without coding covers the complete stack; to frame enterprise compliance, the local LLM and GDPR guide goes through the AI Act obligations point by point.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.