Best open-source LLM for education and teaching
Choose one Open-source education LLM comes down to balancing the instructional quality of the reasoning, the ability to explain a concept step by step, and the hardware efficiency required by deployment in an institution or on a teacher's workstation. The Open-source education LLM offers a decisive advantage over proprietary services: control over student data, no quotas, simpler GDPR compliance, and zero marginal cost after installation. This article compares the most relevant open models for tutoring, exercise generation, homework help, automated grading, and multilingual educational content production, along with their VRAM requirements, licenses, and concrete use cases for moving from dependence on Khanmigo to a sovereign local tutor.
What criteria should an educational LLM meet?
A model intended for education isn't chosen the same way as a coding model. The priority areas are:
- Multi-step reasoning : ability to break down a math or physics problem without skipping logical links. The AIME and MATH scores (published on Papers With Code) are useful indicators.
- General knowledge (MMLU) : subject coverage (history, biology, economics). Models scoring > 80% on MMLU are strong candidates for secondary and higher education.
- French-language quality : critical for French-speaking audiences. European models such as Mistral Large 3 675B or Apertus 70B (Swiss AI) have a native advantage, as do Salamandra 40B Instruct trained on multilingual European corpora.
- Permissive license : Apache 2.0 or MIT to allow deployment in academia without legal friction. The Llama Community license imposes restrictions on the number of monthly active users.
- Context window : to ingest a manual or several copies. Llama 4 Scout 109B announces 10 million context tokens, Seed-OSS 36B Instruct offers 524 288 tokens.
- Controlled hallucinations : a tutor that makes up a formula is worse than a textbook. Prefer models with an explicit "reasoning" mode such as DeepSeek R1 671B or QwQ 32B.
For a broader analysis of the selection criteria, see the general guide to choosing an LLM on quelllm.fr.
Academic tutoring and middle/high school: 30–70B models
For a teaching workstation equipped with a 24 GB card (RTX 4090, RTX 5090) or a dual-GPU workstation, the 30–70B class is the sweet spot. It delivers satisfactory reasoning without saturating VRAM.
- Qwen 3 32B (Apache 2.0): ~19 GB Q4 VRAM, 131 072-token context. Excellent at mathematics and written French, with native chain-of-thought support. Detailed profile on HuggingFace Qwen.
- Gemma 4 31B (license Gemma): 18 GB in Q4, 256,000-token context. Well-calibrated responses for a school audience, with a measured tone.
- Llama 3.3 70B Instruct (Llama 3.3 Community): 40 GB in Q4. The absolute benchmark for general knowledge, with a high MMLU score. Avoid it if the institution exceeds 700 M active monthly users — a threshold rarely reached in education.
- DeepSeek R2 32B (MIT): 19 GB in Q4, explicit reasoning. A good candidate for help with math and physics-chemistry exercises.
- OLMo 3 32B (Apache 2.0, Allen AI): 19 GB in Q4. Distinctive feature: a completely open model (weights + data + training recipe), valuable for academic use where reproducibility matters. See the official Allen AI repo.
- Apertus 70B (Apache 2.0): 40 GB in Q4. European multilingual model designed by Swiss AI with a focus on sovereignty.
For a quantitative comparison, see the page Qwen 3 32B vs Llama 3.3 70B.
Frugality: 30–40B models accessible to a consumer GPU
For a high school equipped with a single RTX 4090, or a shared resource center sharing one workstation, targeting Q4 VRAM under 24 GB remains necessary.
- Qwen 3.6 35B-A3B (Apache 2.0): 21 GB in Q4, MoE architecture with 3B active parameters. High inference speed because only the relevant experts are activated for each token. Ideal for a responsive classroom tutor.
- Seed-OSS 36B Instruct (Apache 2.0, ByteDance): 22 GB in Q4, context 524 288 tokens. Lets you analyze a complete manual or a long essay without splitting it up.
- Salamandra 40B Instruct (Apache 2.0, Barcelona Supercomputing Center): 24 GB in Q4. Trained specifically for European languages, including French and Spanish. Documented on HuggingFace BSC-LT.
- Yi 1.5 34B Chat (Apache 2.0, 01.AI): 20 GB in Q4. Good versatility, but context is limited to 4,096 tokens — best reserved for short interactions.
Estimated tokens/sec on RTX 4090 in Q4: 35–50 tok/s for dense 32B models, 80–120 tok/s for MoE models such as Qwen 3.6 35B-A3B. See the speed benchmark on quelllm.fr.
Use case: generate exercises, grade assignments, explain a concept
Generating differentiated exercises. A teacher provides a learning objective and three levels (support, standard, advanced). Mistral Small 4 (Apache 2.0, 119B, 72 GB in Q4) produces well-formed statements in French and follows the difficulty requirements. Qwen 3.5 122B-A10B (73 GB in Q4) with its 10B active parameters offers an excellent speed/quality ratio for mass production.
Short-answer grading. The function-grid (prompt + grading rubric + student submission) requires rigor. DeepSeek R1 Distill 32B (MIT, 19 GB in Q4) or QwQ 32B (Apache 2.0) handle structured rubrics well. Keeping a human teacher in the loop remains necessary for final grading.
Khanmigo-style conversational tutoring. The goal is to guide without giving the answer. Llama 3.3 70B Instruct et Gemma 4 31B have behavior aligned with Socratic pedagogy. For a Khanmigo alternative entirely local, these two models combined with document-based RAG over official curricula provide a solid foundation.
Multimodal explainers. To explain an SVT diagram or comment on a historical figure, switch to a vision-language model: Qwen 3 VL 235B-A22B (142 GB in Q4) or LLaVA-OneVision 72B (42 GB in Q4) for properly equipped institutions.
Sovereignty, compliance, and on-premises hosting
Deploying a local AI tutor in a school setting raises three specific questions:
- Hosting : an academic or municipal server, with no outbound traffic to a foreign cloud. The MIT and Apache 2.0 models (DeepSeek R2 32B, Mistral Small 4, OLMo 3 32B) do not restrict internal redistribution.
- GDPR compliance : no student prompt is exposed to a third party. The European regulation (EUR-Lex 2016/679) is easier to meet with an on-premises model.
- Filtering and security : an upstream moderation system (classifier, prompt blacklist) remains necessary. No open LLM is inherently safe for minors without explicit safeguards.
To explore the implications of a sovereign architecture, see the on-premises deployment guide on quelllm.fr.
FAQ
Q: Which open-source LLM should you choose for a single-GPU budget RTX 4090?
On a single RTX 4090 (24 GB VRAM), target Qwen 3 32B or Gemma 4 31B in Q4 quantization (~19 GB) leaves room for context. For a better speed/quality ratio, Qwen 3.6 35B-A3B in MoE delivers faster responses thanks to its 3B active parameters.
Q: Is there a real local alternative to Khanmigo?
Yes. Llama 3.3 70B Instruct or Apertus 70B combined with a RAG document base (manuals, official curricula) and a Socratic educational system prompt reproduce Khanmigo’s tutoring function without transmitting student data to a third party. Plan on ~40 GB of VRAM and a workstation with 2 24 GB GPUs.
Q: Are European models sufficient for academic French?
Mistral Large 3 675B et Mistral Small 4 (Apache 2.0) excellent in French and compliant with typographic conventions. Salamandra 40B Instruct (Apache 2.0, BSC) and Apertus 70B (Swiss AI) are trained on European multilingual corpora with a high proportion of French, making them relevant for sovereign use.
Q: Which license should be preferred for deployment in academia?
Apache 2.0 and MIT are the safest: no commercial-use restrictions and no user threshold. The Llama Community license imposes conditions beyond 700 million monthly active users—rarely a constraint for an institution, but confirm with the academy's legal department.
Q: Do you need a “reasoning” model for math tutoring?
For middle school, a general-purpose model such as Gemma 4 31B is sufficient. For science-focused high school and higher education, switch to DeepSeek R1 Distill 32B, QwQ 32B or DeepSeek R2 32B that expose their chain of thought. The AIME scores published on HuggingFace Open LLM Leaderboard confirm their superiority on multi-step problems.
Q: How can you avoid hallucinations in an educational setting?
Three measures that can be combined: (1) connect the model to a RAG database of official manuals and curricula, (2) systematically display the chain of thought when a reasoning model is used, (3) constrain the system prompt ("if you're not sure, say so explicitly"). No single measure eliminates errors.
Conclusion
Choosing a Open-source education LLM depends on the available hardware, target language, and teaching level. For middle and high school on a single GPU, Qwen 3 32B, Gemma 4 31B or Qwen 3.6 35B-A3B are solid entry points; for sovereign deployment at an institutional scale, Mistral Small 4, Apertus 70B or Llama 3.3 70B Instruct offer quality comparable to proprietary services. To fine-tune based on your exact hardware, use the quelllm.fr configurator or browse the complete catalog of the 249 models.