CroissantLLM and Lucie: the 100% French models in local
CroissantLLM, Lucie, Pleias, Mistral: the French open-weight LLM scene is very real, and several of these models run comfortably locally on a mainstream computer. This guide surveys the 2026 landscape of 100% French, self-hostable LLMs, compares their actual French-language quality with the best international models (Qwen 3.8, Gemma 4), and explains when the sovereignty argument justifies choosing Lucie over a Chinese or American model.
#Why a 100% French model?
Most high-performing open-weight LLMs in 2026 come from China (Qwen, DeepSeek, GLM) or the United States (Llama, Gemma, Phi). They often speak French very well — Qwen 3.8 and Gemma 4 are excellent, incidentally — but they were trained primarily on English and Chinese content, with cultural biases that show through: implicit references, awkwardly translated American phrasing, political perspectives, and an imperfect handling of formal and informal language.
A model trained in France addresses three distinct needs: native quality in formal French, culturally relevant content (law, administration, literature, media), and legal sovereignty in the sense that the weights, datasets, and developers fall under European law. For a law firm, government agency, or public research project, this is often decisive.
#Overview of French LLMs in 2026
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Five families dominate today’s self-hostable French landscape. They are not in the same league in terms of size, quality, or licensing.
- Mistral (Mistral AI, Paris)
- The champion. Mistral Small 24B, Magistral, Devstral. Apache 2.0 for the open models, 128k context, French quality on par with the best cloud models.
- CroissantLLM (Centrale Supélec / Illuin)
- A truly open-source 1.3B bilingual FR-EN model: weights, datasets, and code. Designed for research and edge deployment.
- Lucie (LINAGORA / OpenLLM-France)
- 7B Apache 2.0, community project, trained on French public supercomputers (Jean Zay). The first true French model of this size.
- Pleias (Paris)
- Family of 1B to 3B models trained exclusively on open data (Common Corpus). Strict compliance, perfect for the public sector.
- Croco / CroissantCool / derivatives
- Community fine-tunes for French instructions. Often based on Mistral or CroissantLLM, hosted on Hugging Face.
Alongside them, some non-French models stand out in French: Qwen 3.8 (Alibaba), Gemma 4 (Google), and gpt-oss (OpenAI). We put them into perspective below in the benchmarks.
#CroissantLLM: 1.3B bilingual, fully open
CroissantLLM was created through a collaboration between CentraleSupélec, Illuin Technology, Inria, and INRIA Paris. It is a 1.3-billion-parameter model trained on approximately 3 trillion tokens, split evenly between French and English. Small, certainly, but its distinctive feature is complete transparency: weights, datasets, training recipes, evaluation scripts—everything is published.
- Size
- 1.3B parameters, ~800 MB in Q4. Runs on a CPU, a recent smartphone, or a Raspberry Pi 5.
- Languages
- Balanced FR-EN bilingual (1:1 in training).
- License
- MIT, commercial use permitted without restriction.
- Variants
- CroissantLLMBase (pretrained), CroissantLLMChat (instruction-tuned), CroissantCool (community RLHF).
- Usage
- Embedded, edge, research, local fine-tuning. Not a general-purpose assistant to pit against GPT-5.
#Lucie: the French open-source bet
Lucie is LINAGORA's flagship project, developed in partnership with OpenLLM-France and several public research labs. Introduced in early 2025, it is a 7-billion-parameter model trained on the Jean Zay supercomputer at CNRS-IDRIS, on a corpus combining French, English, German, Spanish, Italian, and code, with strong weighting toward French.
- Size
- 7B parameters, ~4.5 GB in Q4_K_M, ~5.5 GB in Q5_K_M.
- VRAM
- 5 GB minimum in Q4. Runs comfortably on RTX 3060 12 GB, or on a 16 GB unified-memory Mac with an M-series chip.
- License
- Apache 2.0, so it can be used commercially and redistributed.
- Context
- 32k tokens in chat version (Lucie-7B-Instruct).
- Sovereignty
- Weights hosted on HuggingFace, but the code and data are accessible. Self-hosted mirroring is possible for government agencies.
What stands out in use: Lucie handles French very cleanly, without the “English calque” still noticeable in Llama 4 with literary phrasing. It is, however, a notch below the best models in its tier, such as Qwen 3.5 9B, for code and mathematical reasoning.
#Pleias: radical compliance through data
Pleias takes a rare approach: it trains its models only on Common Corpus—a public, traceable dataset with no copyrighted content. The result is 1B and 3B models that an administration can adopt without concern about intellectual-property issues, but whose capabilities remain modest compared with large models trained on the entire web.
- Sizes
- Pleias-1B and Pleias-3B (Pico, Nano versions). RAG-friendly.
- Specialty
- Good at RAG, classification, and structured extraction tasks. Weak at free-form chat.
- Argument
- No training on copyright-protected data. Supply chain auditing possible.
- Typical use case
- Summarizing public documents, processing administrative requests, and classifying mail.
Pleias is not a conversational assistant; it is an extraction and structuring tool. For a public-sector organization subject to the AI Act and a strict accountability chain, it is probably the only French model that meets all the requirements without compromise.
#Mistral: the French reference
No need to go longer: Mistral remains the French open-weights leader. Mistral Small 24B (Apache 2.0, ollama run mistral-small, 14 GB in Q4) is probably the best self-hostable French model in 2026 on enthusiast hardware. For code, Devstral 24B (Apache 2.0) is the in-house agent specialist. Mistral Magistral brings reasoning (chain of thought) in French, covered in a dedicated guide on the site.
- Mistral Small 24B
- ~14 GB in Q4 (ollama run mistral-small). RTX 4090 or MacBook Pro M3/M4 with 32+ GB. Probably the best French/quality compromise.
- Devstral 24B
- ~14 GB in Q4, Apache 2.0 (ollama run devstral:24b). The code-agent specialist from Mistral AI.
- Mistral Magistral 24B
- Chain-of-thought reasoning in French.
- Codestral 22B
- Code-specialized but non-production license: prohibited for work. Prefer Devstral 24B.
#FR quality tests vs. Qwen 3.8 and Gemma 4
Here is a qualitative comparison grid based on typical uses by a French-speaking user (formal writing, PDF summarization, general legal questions, coding). Observations were conducted at temperature 0.3, 8k context, Q4_K_M, with identical prompts.
- Mistral Small 24B (FR)
- Excellent writing, polished register, rare agreement errors. Very good code. French reference.
- gpt-oss 20B (US)
- Excellent multilingual performance, with French exceeding expectations. Some English layers in literary text. Very fast (MXFP4).
- Qwen 3.8 27B (CN)
- Very good formal French, sometimes awkward in casual language. Code at Mistral's level, superior reasoning—consider setting reasoning to low, or it will overthink.
- Lucie 7B (FR)
- Clean French, with a fluent formal administrative register. Average coding ability. Good for writing and summarization.
- Qwen 3.5 9B (CN)
- Strong in French for its size (8 GB), 256k context, vision. A good writing and coding compromise on a small configuration.
- Gemma 4 12B (US)
- Good French, sometimes overly academic. Multimodal, good RAG, Apache 2.0.
- CroissantLLM 1.3B (FR)
- The level of a lightweight assistant; not in the same category. Excellent for its size.
#Sovereignty and the AI Act
The European AI Act, which came into force in 2024–2026, imposes several obligations on providers and deployers of general-purpose AI models: technical documentation, a summary of training data, labeling of generated content, and systemic-risk assessments for very large models. Models trained and published in the EU simplify auditing, but the AI Act does not say that you need a French model—it says you must be able to document and explain the system.
- Local hosting argument
- Regardless of the model, running it locally resolves 90% of GDPR concerns. This matters more than the model's nationality.
- Transparency argument
- For an AI Act audit, a model whose datasets are public (Pleias, CroissantLLM, Lucie) is easier to defend than a model whose training is opaque.
- Ecosystem argument
- Supporting Mistral, LINAGORA, and Pleias helps consolidate a European industry. It is a political argument, but a real one for public procurement.
- Legal argument
- The model weights published under Apache 2.0 by a European entity reduce uncertainty around data transfers to the US.
#Which one to choose for your use case
- 01You're just getting started and want to test a French modelInstall Lucie 7B (ollama pull OpenLLM-France/Lucie-7B-Instruct-gguf:Q4_K_M), ~5 GB, a truly 100% French model that runs everywhere. For a stronger general-purpose French model, run ollama run mistral-small (Mistral Small 24B, 14 GB).
- 02You have 12 GB of VRAM (RTX 3060/4070)Lucie-7B-Instruct in Q5_K_M for pure French, or Qwen 3.5 9B in Q8 (11 GB) for a more capable generalist. You can also try Mistral Small 24B in Q3.
- 03You have 24 GB (RTX 4090/3090, Mac 32+ GB)Mistral Small 24B in Q4_K_M is probably the best self-hostable French LLM, period. Add Lucie 7B for fast “pure French” use cases.
- 04You are a government agency or public-sector organizationPleias 1B/3B for document RAG (zero copyright risk), Lucie 7B for general assistance, Mistral Small 24B when you need quality. Avoid models with unclear licensing.
- 05You use embedded / edge / Raspberry PiCroissantLLM 1.3B Q4. It is the smallest French model that remains usable. See the LLM guide for Raspberry Pi for details.
#Go further
Three related paths depending on your next step: to go deeper on French reasoning, the guide dedicated to Mistral Magistral details the FR chain of thought; to build a RAG system over your PDFs with one of these models, the local RAG guide with Ollama without coding covers the complete stack; to frame enterprise compliance, the local LLM and GDPR guide goes through the AI Act obligations point by point.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.