Tesseract OCR: read a scan locally (guide French)
Tesseract is the leading open-source OCR engine (Apache 2.0, more than 70,000 GitHub stars, version 5.5.3 in 2026): free, offline, available everywhere, and its accuracy depends far more on image quality—resolution, deskewing, and contrast—than on its own settings. On clean text, it remains fast and reliable; with a complex layout or a slightly tilted document, its accuracy can suddenly collapse.
A scanned PDF contains no text: it is an image of text. Until it has gone through optical character recognition, it indexes as empty, and the document assistant insists that the document says nothing. Tesseract is the standard open-source engine for this step: free, offline, available everywhere, and far more dependent on image quality than on its own settings.
#What Tesseract is—and isn't
Tesseract is an optical character recognition engine released under the Apache 2.0 license. The official repository, described by its maintainers as the “Tesseract Open Source OCR Engine,” now exceeds 70,000 stars on GitHub. Version 4 added an engine based on LSTM neural networks, focused on line recognition; the historical engine, which recognizes character patterns, remains available as an option. It reads common image formats (PNG, JPEG, TIFF) and produces plain text, hOCR, PDF, or TSV. The latest stable version as of this writing, 5.5.3, was released on July 24, 2026. In practice: install the language file, prepare the image carefully (300 dots per inch, straight page, clear contrast), choose the correct segmentation mode, then check the result on about twenty pages before launching a batch.
What it is not: a document-understanding tool. It outputs text, possibly with positions. It does not identify a block as a heading, reconstruct a table into rows and columns, or understand what it reads. For a structured document, OCR is just one component among several.
#French: the language file
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
By default, Tesseract reads English. On a French document, this results in lost accents and approximate words—and many people conclude that the engine is poor. The documentation specifies that without the -l option, English is assumed. You therefore need to install the language data file and request it explicitly at runtime. The project claims recognition of more than one hundred languages once the correct packages are installed.
#Image quality is everything
This is the point people discover too late: the official documentation devotes its “Improve quality” page to image processing before engine settings. Tesseract already performs internal operations through the Leptonica library, but they are not always sufficient, and a few lines of preprocessing are often enough to make a mediocre scan readable. Koncile, an invoice-extraction software vendor, illustrates the scale of the issue in its own tests, published in 2026: for an invoice tilted by barely three to five degrees, the correct reading rate drops from 100% to 31%. This is not an independent study, but the order of magnitude matches the warning in the official documentation. No engine setting can recover from a drop that severe; only prior deskewing can.
- Resolution
- The official documentation states that Tesseract works best on images of at least 300 dots per inch and suggests resizing them otherwise.
- Reframing
- The documentation warns that line segmentation quality drops significantly when the page is too tilted: you need to rotate the image so the lines are horizontal. See the figure above.
- Contrast and binarization
- Tesseract binarizes internally (Otsu's algorithm), which can produce mediocre results when the background has an uneven shade. Version 5.0 added two methods, Adaptive Otsu and Sauvola, configurable through the thresholding_* parameters. For an engine running version 4 or later, you need dark text on a light background.
- Borders
- Very wide borders confuse the engine, but tightly cropped text does too: the documentation recommends adding a small white border (10 pixels in its example) around a borderless excerpt.
- 01Straighten the pageDetect the text's dominant angle and rotate the image before anything else; a simple script based on the Hough transform is sufficient for most office scanners.
- 02Clean up the contrastConvert to grayscale, then binarize (use adaptive thresholding rather than a fixed threshold, which fails under uneven lighting) to produce solid black on solid white.
- 03Crop and run OCRRemove scanner edges and large blank margins, leave a thin white border around the text, then call Tesseract with the language and segmentation mode suited to the document type being processed.
#Settings that change the result
| Option | What it is used for | When to touch it |
|---|---|---|
| Page segmentation mode (--psm) | Tells the engine what the page looks like: single block, column, isolated line (7), scattered text (11) | On a ticket, label, narrow column, or small crop: the default mode expects an entire page |
| Engine mode (--oem) | 1: LSTM neural network only; 0: historical engine | Rarely: the language files in common Linux packages (tessdata_fast) support only LSTM, and modes 0 and 2 do not work with them |
| List of allowed characters (tessedit_char_whitelist) | Restricts recognition to a subset of characters | For a purely numeric field, to limit confusion between letters and numbers |
| Dictionaries (load_system_dawg, load_freq_dawg) | Two variables that enable the engine’s word lists | On receipts, price lists, and codes: disabling them can help when most of the text is not made up of common words |
Segmentation mode is the setting people forget most often. The documentation makes this clear: by default, Tesseract expects a page of text, and you need to choose a different mode to read a small area. On an image containing a single line, mode 7 treats it as one line. For sparse text with no particular order—such as a screenshot with labels scattered around—the corresponding mode avoids the same problem. The engine mode, on the other hand, rarely changes the result, for a practical reason: the legacy engine is present only in the language files from the tessdata repository, not in the fast or accurate variants shipped by most distributions.
On clean, well-framed printed text, Tesseract remains competitive: FastOCR, an online OCR service selling a cloud alternative, reports 95 to 97.2% accuracy for Tesseract v5 on clean English text, versus 98.2 to 99.1% for Google Cloud Vision. These are figures from an interested party, with no published protocol: treat them as an order of magnitude, not a measurement. The engine should be chosen based on the actual nature of the documents: for clean, printed documents, Tesseract is more than sufficient; for damaged, handwritten, or densely structured documents, it becomes the weakest link in the chain.
#Tesseract or a vision model
| Criterion | Tesseract | Local vision model |
|---|---|---|
| Resources | CPU only | GPU recommended, several gigabytes |
| Speed | Very fast | Significantly slower |
| Clean text in one column | Excellent | Equivalent, more expensive |
| Complex layouts, tables | Low: the project documentation acknowledges a known issue with arrays | Much better |
| Handwriting | Very low: 25 to 40% according to FastOCR | Significantly better |
| Content understanding | None | Can answer a question about the document |
| Risk of fabrication | Weak: it confuses characters rather than composing values | In reality: a model can hallucinate a plausible value |
This last point matters for any document review: Tesseract can mistake one character for another, while a vision model can produce a perfectly formatted but incorrect number. The distinction isn't absolute, because a 0 misread as an 8 in an amount won't be obvious either, but nonsensical text more often gives away an OCR engine than a plausible invented value. A comparison published in 2026 sums up the changed landscape: the real alternatives are no longer conventional OCR engines but open-source vision-language models, explicitly including PaddleOCR-VL—at a cost, since these models require a GPU where Tesseract is fine with an ordinary processor.
- Local multimodal models and what they can read
- Extract data from an invoice: the complete pipeline
- Convert structured documents with Docling
- PaddleOCR: the engine for handling complex layouts
- Source: official Tesseract repository on GitHub
- Source: Koncile, 2026 Tesseract comparison (invoice extraction editor)
- Source: tesseract manual page (the --psm and --oem options)
- Source: Tesseract documentation, improving quality
- Source: Tesseract documentation, data files
#The true cost of a Tesseract project
The engine itself costs nothing, which often obscures where an OCR project's time actually goes. The real work lies around the engine: image preprocessing (deskewing, binarization, borders) and result verification. A realistic budget sets aside time for both steps from the outset rather than discovering them after a first batch of disappointing results.
- Phone scanner or photo
- A flatbed scanner produces straighter, better-lit pages than a handheld photo; use a scanner whenever the volume justifies it.
- Output format
- Plain text is enough for indexing; the hOCR format preserves positions, which is useful for highlighting an excerpt in the original interface.
- Quality control
- Manually reviewing one sample after each change of scan source catches drift before it affects the entire batch.
- Volume and parallelization
- Tesseract is multithreaded by default (OpenMP), which is unsuitable for a large batch. The project FAQ recommends launching one process per core with the OMP_THREAD_LIMIT=1 variable, which eliminates multithreading overhead.
One often-overlooked final item: naming the files produced. Across several thousand pages, you need to be able to identify which text corresponds to which scan; keeping the same base name for the image and the recognized text solves the problem from the start.
#Insert it into a local pipeline
In a local document-processing setup, Tesseract is used in only one place: just before indexing, on documents without a text layer. The test to automate is simple—if text extraction from a PDF returns fewer than a few dozen characters per page (a threshold to adjust for your corpus), it is probably a scan and should be sent to OCR. Without this test, entire documents enter the index silently empty, and no one notices until a user is surprised.
Deskewing and binarization should be automated in this same pipeline rather than left to the judgment of each page: a script that detects the dominant angle and corrects it before calling the engine avoids many of the unpleasant surprises described above. On a homogeneous batch—the same scanner and the same document type every time—you configure this preprocessing once and it then runs unattended.
#FAQ
Is Tesseract free?+
How do you make it read French?+
Why is my output unreadable?+
Can Tesseract read tables?+
Should you choose a vision model?+
What is page segmentation mode?+
How do you make a scanned PDF searchable with Tesseract?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.