Translating documents locally: the alternative to DeepL
Pasting a contract, financial report, or HR file into DeepL or ChatGPT means sending that document to servers you do not control. Local AI translation solves this problem: a multilingual LLM runs on your machine, and no data leaves it. This guide shows which models to choose, how to preserve file formatting, and how to translate dozens of documents in batches with a simple Ollama script.
#Why not send your documents to DeepL
DeepL and Google Translate are excellent, but their business model is cloud-based: your text passes through their servers and, depending on the plan, may be retained or used to improve the models. For a harmless email, it doesn't matter. For an NDA-covered contract, a patent application in progress, financial statements, or employees' personal data, it's a real legal and competitive risk.
Local AI translation changes the game: the model is downloaded to your disk and runs on your CPU or GPU. No network requests, no account, no volume limit. You can translate an entire folder of confidential documents without a single byte leaving your machine—including offline, on a plane, or on an isolated network.
- Privacy
- Contracts, patents, medical data, and HR data stay on your machine.
- Zero usage cost
- No subscription or per-character billing: translate 10 or 10,000 pages for the same price.
- Hors-ligne
- Works offline, useful on an isolated workstation or while traveling.
- Unlimited volume
- No monthly quota or throttling like with cloud APIs.
#Which local models actually translate well
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
Not all LLMs are equal at translation. A good local translator must be trained on a large multilingual corpus and handle French with precision. In 2026, a few families stand out clearly for self-hosted use.
- Qwen 3.5 9B
- Excellent for multilingual use, very good in French, and handles nuances and technical vocabulary well. ≈6.6 GB in Q4, 256k context, Apache 2.0 license: the best quality-to-resource tradeoff in 2026. Move up to Qwen 3.8 27B (≈18 GB) if the quality is not sufficient.
- Gemma 4 12B
- Fluid, natural translations with polished French punctuation. Multimodal, now under the Apache 2.0 license (Gemma switched to Apache 2.0 in 2026). ≈7.6 GB in Q4, comfortable on a 12-16 GB GPU or a Mac with unified memory.
- Qwen 3.5 4B
- The default small multilingual model in 2026: only ≈3.4 GB, it runs everywhere—even on a CPU—and already translates remarkably well for its size. Apache 2.0.
- Mistral Small 24B
- Excellent in French, fast, ≈14 GB in Q4. A solid general-purpose model and a very good choice for processing large volumes on a 16 GB GPU.
- Qwen 3.6 35B-A3B
- Cloud-like quality thanks to its MoE architecture, now accessible with as little as 32 GB (≈23 GB in Q4, RTX 4090 or a Mac with a large amount of unified memory). Reserved for demanding translations.
Specialized “translation” models such as NLLB and Opus-MT also exist and are very lightweight, but a recent general-purpose LLM often outperforms them on long texts and context adherence (register, professional terminology, and consistency across an entire document).
#Requirements and installation
The foundation is conventional: Ollama as the inference engine, a multilingual model, and Python for file-processing scripts. Ollama listens on http://localhost:11434 by default.
- Ollama installed
- The daemon that downloads and runs the models. See the installation guide if you haven't already.
- A comfortable GPU
- A RTX 3060 12 GB runs a Qwen 3.5 9B in Q4 (~6.6 GB) effortlessly. It also works on a CPU, but more slowly.
- Python 3.10+
- To work with DOCX, PDF, and Markdown and call the Ollama API.
#First translation test
Before automating, validate quality on a representative excerpt. The key to good local translation is the system prompt: it defines the role, target language, and above all the instruction to add nothing.
#Preserve formatting (DOCX, PDF, Markdown)
Translating plain text is easy; preserving the formatting of a real document is harder. The right approach is not to feed the entire file to the model, but to translate each text segment in place while leaving the structure intact.
#Word files (DOCX)
python-docx exposes every paragraph and table cell. We translate the text in each “run” while preserving its style, then rewrite the file. The layout, headings, and tables are preserved.
#Markdown
Markdown is the simplest format to translate cleanly: just tell the model to preserve the syntax. Headings, lists, links, and code blocks remain intact if the prompt requires it.
PDF is the trickiest case: it is a presentation format, not a content format. Extract the text with pypdf, translate it, then output it in a new document (often a DOCX or a generated PDF). Reproducing a PDF's exact layout is rarely worthwhile; aim for a readable document rather than a pixel-perfect copy.
#Batch translation with Ollama
The major benefit of running locally is volume without a bill. A script that scans a folder and translates every file turns your workstation into a batch translation service. Here’s a script that processes all .docx files in a directory.
#Improve translation quality
A local LLM translates well by default, but a few adjustments take the result from “acceptable” to “publishable,” especially for business texts.
- 01Provide a glossaryAdd a glossary of business terms and their required translations to the system prompt (legal entity names, product names, acronyms). The model will follow your terminology instead of making things up.
- 02Split intelligentlyTranslate by paragraph or section, not sentence by sentence: the model needs context to handle agreement, pronoun references, and register properly.
- 03Provide the document's contextSpecify the nature of the text (“business contract,” “technical documentation,” “internal email”) in the prompt. The register and vocabulary will adjust accordingly.
- 04Review critical passagesFor a contract or legal document, human review of sensitive clauses remains essential—the local setup gives you privacy, not infallibility.
- 05Move up to a larger model if neededIf quality plateaus, move from Qwen 3.5 9B to Qwen 3.8 27B (set its reasoning to « low »: by default it overthinks, which is unnecessary for translation), or even to an MoE such as Qwen 3.6 35B-A3B. The improvement is clear on long sentences and fine-grained terminology.
#Troubleshooting
- The model adds comments
- Strengthen the instruction “return ONLY the translation” and lower the temperature. Some models prefix it with “Here is the translation:”—remove that in post-processing if needed.
- It translates into the wrong language
- Explicitly name the target language in the prompt and provide an example. With a short text, a model may choose the wrong language.
- DOCX formatting is lost
- This is the replacement for para.text that flattens runs. Translate run by run for heavily formatted documents.
- Slow translation for large volumes
- Make sure the model fits in VRAM (otherwise CPU offload means slow performance). Reduce the model size or use OLLAMA_KEEP_ALIVE.
- Translated Markdown code blocks
- Emphasize it in the prompt: "don't translate the contents of code blocks." Otherwise, extract them before translation and reinsert them afterward.
- Context too long and truncated
- A very large document exceeds the context window. Split it into sections and translate each separately.
#Go further
Local AI translation relies on a well-chosen Ollama foundation. These site guides complete the setup:
- Install Ollama
- To set up the inference engine if you haven't already.
- Q4, Q5, Q8: which quantization should you choose
- To balance translation quality, speed, and VRAM for your GPU.
- n8n + Ollama: automate with a local AI
- To connect translation to a no-code workflow (incoming emails, dropped-off files, etc.).
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.