Intermediate 11 minTranslation

Translating documents locally: the alternative to DeepL

Pasting a contract, financial report, or HR file into DeepL or ChatGPT means sending that document to servers you do not control. Local AI translation solves this problem: a multilingual LLM runs on your machine, and no data leaves it. This guide shows which models to choose, how to preserve file formatting, and how to translate dozens of documents in batches with a simple Ollama script.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why not send your documents to DeepL

DeepL and Google Translate are excellent, but their business model is cloud-based: your text passes through their servers and, depending on the plan, may be retained or used to improve the models. For a harmless email, it doesn't matter. For an NDA-covered contract, a patent application in progress, financial statements, or employees' personal data, it's a real legal and competitive risk.

Local AI translation changes the game: the model is downloaded to your disk and runs on your CPU or GPU. No network requests, no account, no volume limit. You can translate an entire folder of confidential documents without a single byte leaving your machine—including offline, on a plane, or on an isolated network.

i
The real argument: compliance
For a law firm, an HR department, or a healthcare team, local translation eliminates any GDPR concerns about transferring data to a processor. The document never leaves the company’s perimeter.
Privacy
Contracts, patents, medical data, and HR data stay on your machine.
Zero usage cost
No subscription or per-character billing: translate 10 or 10,000 pages for the same price.
Hors-ligne
Works offline, useful on an isolated workstation or while traveling.
Unlimited volume
No monthly quota or throttling like with cloud APIs.

#Which local models actually translate well

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Not all LLMs are equal at translation. A good local translator must be trained on a large multilingual corpus and handle French with precision. In 2026, a few families stand out clearly for self-hosted use.

Qwen 3.5 9B
Excellent for multilingual use, very good in French, and handles nuances and technical vocabulary well. ≈6.6 GB in Q4, 256k context, Apache 2.0 license: the best quality-to-resource tradeoff in 2026. Move up to Qwen 3.8 27B (≈18 GB) if the quality is not sufficient.
Gemma 4 12B
Fluid, natural translations with polished French punctuation. Multimodal, now under the Apache 2.0 license (Gemma switched to Apache 2.0 in 2026). ≈7.6 GB in Q4, comfortable on a 12-16 GB GPU or a Mac with unified memory.
Qwen 3.5 4B
The default small multilingual model in 2026: only ≈3.4 GB, it runs everywhere—even on a CPU—and already translates remarkably well for its size. Apache 2.0.
Mistral Small 24B
Excellent in French, fast, ≈14 GB in Q4. A solid general-purpose model and a very good choice for processing large volumes on a 16 GB GPU.
Qwen 3.6 35B-A3B
Cloud-like quality thanks to its MoE architecture, now accessible with as little as 32 GB (≈23 GB in Q4, RTX 4090 or a Mac with a large amount of unified memory). Reserved for demanding translations.
→
Where to start
If you're unsure, run Qwen 3.5 9B in Q4_K_M. It offers the best quality/VRAM ratio for FR ↔ EN/ES/DE/IT translation on an 8–12 GB GPU. Switch to Qwen 3.8 27B only if the quality is not sufficient.

Specialized “translation” models such as NLLB and Opus-MT also exist and are very lightweight, but a recent general-purpose LLM often outperforms them on long texts and context adherence (register, professional terminology, and consistency across an entire document).


#Requirements and installation

The foundation is conventional: Ollama as the inference engine, a multilingual model, and Python for file-processing scripts. Ollama listens on http://localhost:11434 by default.

Ollama installed
The daemon that downloads and runs the models. See the installation guide if you haven't already.
A comfortable GPU
A RTX 3060 12 GB runs a Qwen 3.5 9B in Q4 (~6.6 GB) effortlessly. It also works on a CPU, but more slowly.
Python 3.10+
To work with DOCX, PDF, and Markdown and call the Ollama API.
Terminal
# Récupérer un modèle multilingue adapté à la traduction
ollama pull qwen3.5:9b

# Vérifier qu'il répond
ollama run qwen3.5:9b "Traduis en anglais : Le contrat prend effet à sa signature."
Terminal
# Dépendances Python pour les formats de documents
pip install ollama python-docx pypdf markdown-it-py

#First translation test

Before automating, validate quality on a representative excerpt. The key to good local translation is the system prompt: it defines the role, target language, and above all the instruction to add nothing.

traduire.py
import ollama

MODELE = "qwen3.5:9b"

SYSTEME = (
    "Tu es un traducteur professionnel. Traduis le texte fourni de "
    "{source} vers {cible}. Conserve le sens, le registre et la "
    "terminologie technique. Ne commente pas, n'ajoute rien : "
    "renvoie uniquement la traduction."
)

def traduire(texte, source="français", cible="anglais"):
    reponse = ollama.chat(
        model=MODELE,
        messages=[
            {"role": "system", "content": SYSTEME.format(source=source, cible=cible)},
            {"role": "user", "content": texte},
        ],
        options={"temperature": 0.2},
    )
    return reponse["message"]["content"].strip()

if __name__ == "__main__":
    extrait = "Le présent contrat prend effet à la date de sa signature par les deux parties."
    print(traduire(extrait))
→
Low temperature
For translation, a temperature of 0.1 to 0.3 produces more faithful and reproducible results. A high value encourages the model to “embellish” and rephrase, which is undesirable.

#Preserve formatting (DOCX, PDF, Markdown)

Translating plain text is easy; preserving the formatting of a real document is harder. The right approach is not to feed the entire file to the model, but to translate each text segment in place while leaving the structure intact.

#Word files (DOCX)

python-docx exposes every paragraph and table cell. We translate the text in each “run” while preserving its style, then rewrite the file. The layout, headings, and tables are preserved.

traduire_docx.py
from docx import Document
from traduire import traduire

def traduire_docx(entree, sortie, cible="anglais"):
    doc = Document(entree)

    for para in doc.paragraphs:
        if para.text.strip():
            para.text = traduire(para.text, cible=cible)

    for table in doc.tables:
        for ligne in table.rows:
            for cellule in ligne.cells:
                if cellule.text.strip():
                    cellule.text = traduire(cellule.text, cible=cible)

    doc.save(sortie)
    print(f"Traduit : {sortie}")

if __name__ == "__main__":
    traduire_docx("contrat.docx", "contrat_en.docx")
!
The style may drift across multiple runs
Rewriting para.text replaces the runs and can flatten fine-grained formatting (bold in the middle of a sentence). For flawless rendering in heavily formatted documents, translate run by run instead of the entire paragraph—at the cost of less natural segmentation for the model.

#Markdown

Markdown is the simplest format to translate cleanly: just tell the model to preserve the syntax. Headings, lists, links, and code blocks remain intact if the prompt requires it.

traduire_md.py
import ollama

SYSTEME_MD = (
    "Traduis ce document Markdown en {cible}. Conserve EXACTEMENT la "
    "syntaxe Markdown : titres (#), listes, liens, tableaux et blocs "
    "de code. Ne traduis PAS le contenu des blocs de code. Renvoie "
    "uniquement le Markdown traduit."
)

def traduire_markdown(chemin, sortie, cible="anglais"):
    texte = open(chemin, encoding="utf-8").read()
    reponse = ollama.chat(
        model="qwen3.5:9b",
        messages=[
            {"role": "system", "content": SYSTEME_MD.format(cible=cible)},
            {"role": "user", "content": texte},
        ],
        options={"temperature": 0.2},
    )
    open(sortie, "w", encoding="utf-8").write(reponse["message"]["content"])

traduire_markdown("README.md", "README_en.md")

#PDF

PDF is the trickiest case: it is a presentation format, not a content format. Extract the text with pypdf, translate it, then output it in a new document (often a DOCX or a generated PDF). Reproducing a PDF's exact layout is rarely worthwhile; aim for a readable document rather than a pixel-perfect copy.

traduire_pdf.py
from pypdf import PdfReader
from docx import Document
from traduire import traduire

def pdf_vers_docx_traduit(entree, sortie, cible="anglais"):
    lecteur = PdfReader(entree)
    doc = Document()

    for page in lecteur.pages:
        texte = page.extract_text() or ""
        for paragraphe in texte.split("\n\n"):
            if paragraphe.strip():
                doc.add_paragraph(traduire(paragraphe, cible=cible))

    doc.save(sortie)

pdf_vers_docx_traduit("rapport.pdf", "rapport_en.docx")

#Batch translation with Ollama

The major benefit of running locally is volume without a bill. A script that scans a folder and translates every file turns your workstation into a batch translation service. Here’s a script that processes all .docx files in a directory.

batch_traduction.py
import sys
from pathlib import Path
from traduire_docx import traduire_docx

def traduire_dossier(dossier, cible="anglais", suffixe="_traduit"):
    dossier = Path(dossier)
    fichiers = list(dossier.glob("*.docx"))
    print(f"{len(fichiers)} fichier(s) à traduire\n")

    for i, fichier in enumerate(fichiers, 1):
        if suffixe in fichier.stem:
            continue  # éviter de retraduire une sortie
        sortie = fichier.with_name(f"{fichier.stem}{suffixe}.docx")
        print(f"[{i}/{len(fichiers)}] {fichier.name}")
        traduire_docx(fichier, sortie, cible=cible)

    print("\nTerminé.")

if __name__ == "__main__":
    dossier = sys.argv[1] if len(sys.argv) > 1 else "."
    traduire_dossier(dossier)
Terminal
# Traduire tous les .docx d'un dossier en anglais
python batch_traduction.py ./documents_a_traduire
→
Keep the model loaded
For a large batch, set OLLAMA_KEEP_ALIVE (e.g., export OLLAMA_KEEP_ALIVE=30m) to prevent Ollama from unloading the model between files. Reloading it into VRAM costs several seconds each time.

#Improve translation quality

A local LLM translates well by default, but a few adjustments take the result from “acceptable” to “publishable,” especially for business texts.

  1. 01
    Provide a glossary
    Add a glossary of business terms and their required translations to the system prompt (legal entity names, product names, acronyms). The model will follow your terminology instead of making things up.
  2. 02
    Split intelligently
    Translate by paragraph or section, not sentence by sentence: the model needs context to handle agreement, pronoun references, and register properly.
  3. 03
    Provide the document's context
    Specify the nature of the text (“business contract,” “technical documentation,” “internal email”) in the prompt. The register and vocabulary will adjust accordingly.
  4. 04
    Review critical passages
    For a contract or legal document, human review of sensitive clauses remains essential—the local setup gives you privacy, not infallibility.
  5. 05
    Move up to a larger model if needed
    If quality plateaus, move from Qwen 3.5 9B to Qwen 3.8 27B (set its reasoning to « low »: by default it overthinks, which is unnecessary for translation), or even to an MoE such as Qwen 3.6 35B-A3B. The improvement is clear on long sentences and fine-grained terminology.

#Troubleshooting

The model adds comments
Strengthen the instruction “return ONLY the translation” and lower the temperature. Some models prefix it with “Here is the translation:”—remove that in post-processing if needed.
It translates into the wrong language
Explicitly name the target language in the prompt and provide an example. With a short text, a model may choose the wrong language.
DOCX formatting is lost
This is the replacement for para.text that flattens runs. Translate run by run for heavily formatted documents.
Slow translation for large volumes
Make sure the model fits in VRAM (otherwise CPU offload means slow performance). Reduce the model size or use OLLAMA_KEEP_ALIVE.
Translated Markdown code blocks
Emphasize it in the prompt: "don't translate the contents of code blocks." Otherwise, extract them before translation and reinsert them afterward.
Context too long and truncated
A very large document exceeds the context window. Split it into sections and translate each separately.

#Go further

Local AI translation relies on a well-chosen Ollama foundation. These site guides complete the setup:

Install Ollama
To set up the inference engine if you haven't already.
Q4, Q5, Q8: which quantization should you choose
To balance translation quality, speed, and VRAM for your GPU.
n8n + Ollama: automate with a local AI
To connect translation to a no-code workflow (incoming emails, dropped-off files, etc.).
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.