Intermediate 12 minVision

Tesseract OCR: read a scan locally (guide French)

Direct response

Tesseract is the leading open-source OCR engine (Apache 2.0, more than 70,000 GitHub stars, version 5.5.3 in 2026): free, offline, available everywhere, and its accuracy depends far more on image quality—resolution, deskewing, and contrast—than on its own settings. On clean text, it remains fast and reliable; with a complex layout or a slightly tilted document, its accuracy can suddenly collapse.

A scanned PDF contains no text: it is an image of text. Until it has gone through optical character recognition, it indexes as empty, and the document assistant insists that the document says nothing. Tesseract is the standard open-source engine for this step: free, offline, available everywhere, and far more dependent on image quality than on its own settings.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#What Tesseract is—and isn't

Tesseract is an optical character recognition engine released under the Apache 2.0 license. The official repository, described by its maintainers as the “Tesseract Open Source OCR Engine,” now exceeds 70,000 stars on GitHub. Version 4 added an engine based on LSTM neural networks, focused on line recognition; the historical engine, which recognizes character patterns, remains available as an option. It reads common image formats (PNG, JPEG, TIFF) and produces plain text, hOCR, PDF, or TSV. The latest stable version as of this writing, 5.5.3, was released on July 24, 2026. In practice: install the language file, prepare the image carefully (300 dots per inch, straight page, clear contrast), choose the correct segmentation mode, then check the result on about twenty pages before launching a batch.

What it is not: a document-understanding tool. It outputs text, possibly with positions. It does not identify a block as a heading, reconstruct a table into rows and columns, or understand what it reads. For a structured document, OCR is just one component among several.

#French: the language file

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

By default, Tesseract reads English. On a French document, this results in lost accents and approximate words—and many people conclude that the engine is poor. The documentation specifies that without the -l option, English is assumed. You therefore need to install the language data file and request it explicitly at runtime. The project claims recognition of more than one hundred languages once the correct packages are installed.

Recognition of a French scan
# Debian/Ubuntu : installer les données françaises
sudo apt install tesseract-ocr tesseract-ocr-fra

# macOS (Homebrew) : le moteur puis les langues supplémentaires
brew install tesseract tesseract-lang

# Reconnaître une image en français, sortie texte
tesseract scan.png sortie -l fra

# Plusieurs langues dans le même document
tesseract scan.png sortie -l fra+eng

# Sortie PDF avec couche de texte interrogeable
tesseract scan.png sortie -l fra pdf
i
A PDF is not an image
Tesseract expects images. For an image, appending pdf at the end of the command produces a PDF with a searchable text layer. For a multi-page scanned PDF, a tool such as OCRmyPDF, which relies on Tesseract, rasterizes the pages and then adds the text layer to the document.

#Image quality is everything

This is the point people discover too late: the official documentation devotes its “Improve quality” page to image processing before engine settings. Tesseract already performs internal operations through the Leptonica library, but they are not always sufficient, and a few lines of preprocessing are often enough to make a mediocre scan readable. Koncile, an invoice-extraction software vendor, illustrates the scale of the issue in its own tests, published in 2026: for an invoice tilted by barely three to five degrees, the correct reading rate drops from 100% to 31%. This is not an independent study, but the order of magnitude matches the warning in the official documentation. No engine setting can recover from a drop that severe; only prior deskewing can.

Resolution
The official documentation states that Tesseract works best on images of at least 300 dots per inch and suggests resizing them otherwise.
Reframing
The documentation warns that line segmentation quality drops significantly when the page is too tilted: you need to rotate the image so the lines are horizontal. See the figure above.
Contrast and binarization
Tesseract binarizes internally (Otsu's algorithm), which can produce mediocre results when the background has an uneven shade. Version 5.0 added two methods, Adaptive Otsu and Sauvola, configurable through the thresholding_* parameters. For an engine running version 4 or later, you need dark text on a light background.
Borders
Very wide borders confuse the engine, but tightly cropped text does too: the documentation recommends adding a small white border (10 pixels in its example) around a borderless excerpt.
  1. 01
    Straighten the page
    Detect the text's dominant angle and rotate the image before anything else; a simple script based on the Hough transform is sufficient for most office scanners.
  2. 02
    Clean up the contrast
    Convert to grayscale, then binarize (use adaptive thresholding rather than a fixed threshold, which fails under uneven lighting) to produce solid black on solid white.
  3. 03
    Crop and run OCR
    Remove scanner edges and large blank margins, leave a thin white border around the text, then call Tesseract with the language and segmentation mode suited to the document type being processed.

#Settings that change the result

Options you need to know
OptionWhat it is used forWhen to touch it
Page segmentation mode (--psm)Tells the engine what the page looks like: single block, column, isolated line (7), scattered text (11)On a ticket, label, narrow column, or small crop: the default mode expects an entire page
Engine mode (--oem)1: LSTM neural network only; 0: historical engineRarely: the language files in common Linux packages (tessdata_fast) support only LSTM, and modes 0 and 2 do not work with them
List of allowed characters (tessedit_char_whitelist)Restricts recognition to a subset of charactersFor a purely numeric field, to limit confusion between letters and numbers
Dictionaries (load_system_dawg, load_freq_dawg)Two variables that enable the engine’s word listsOn receipts, price lists, and codes: disabling them can help when most of the text is not made up of common words

Segmentation mode is the setting people forget most often. The documentation makes this clear: by default, Tesseract expects a page of text, and you need to choose a different mode to read a small area. On an image containing a single line, mode 7 treats it as one line. For sparse text with no particular order—such as a screenshot with labels scattered around—the corresponding mode avoids the same problem. The engine mode, on the other hand, rarely changes the result, for a practical reason: the legacy engine is present only in the language files from the tessdata repository, not in the fast or accurate variants shipped by most distributions.

On clean, well-framed printed text, Tesseract remains competitive: FastOCR, an online OCR service selling a cloud alternative, reports 95 to 97.2% accuracy for Tesseract v5 on clean English text, versus 98.2 to 99.1% for Google Cloud Vision. These are figures from an interested party, with no published protocol: treat them as an order of magnitude, not a measurement. The engine should be chosen based on the actual nature of the documents: for clean, printed documents, Tesseract is more than sufficient; for damaged, handwritten, or densely structured documents, it becomes the weakest link in the chain.

#Tesseract or a vision model

Two families, two use cases
CriterionTesseractLocal vision model
ResourcesCPU onlyGPU recommended, several gigabytes
SpeedVery fastSignificantly slower
Clean text in one columnExcellentEquivalent, more expensive
Complex layouts, tablesLow: the project documentation acknowledges a known issue with arraysMuch better
HandwritingVery low: 25 to 40% according to FastOCRSignificantly better
Content understandingNoneCan answer a question about the document
Risk of fabricationWeak: it confuses characters rather than composing valuesIn reality: a model can hallucinate a plausible value

This last point matters for any document review: Tesseract can mistake one character for another, while a vision model can produce a perfectly formatted but incorrect number. The distinction isn't absolute, because a 0 misread as an 8 in an amount won't be obvious either, but nonsensical text more often gives away an OCR engine than a plausible invented value. A comparison published in 2026 sums up the changed landscape: the real alternatives are no longer conventional OCR engines but open-source vision-language models, explicitly including PaddleOCR-VL—at a cost, since these models require a GPU where Tesseract is fine with an ordinary processor.

#The true cost of a Tesseract project

The engine itself costs nothing, which often obscures where an OCR project's time actually goes. The real work lies around the engine: image preprocessing (deskewing, binarization, borders) and result verification. A realistic budget sets aside time for both steps from the outset rather than discovering them after a first batch of disappointing results.

Phone scanner or photo
A flatbed scanner produces straighter, better-lit pages than a handheld photo; use a scanner whenever the volume justifies it.
Output format
Plain text is enough for indexing; the hOCR format preserves positions, which is useful for highlighting an excerpt in the original interface.
Quality control
Manually reviewing one sample after each change of scan source catches drift before it affects the entire batch.
Volume and parallelization
Tesseract is multithreaded by default (OpenMP), which is unsuitable for a large batch. The project FAQ recommends launching one process per core with the OMP_THREAD_LIMIT=1 variable, which eliminates multithreading overhead.

One often-overlooked final item: naming the files produced. Across several thousand pages, you need to be able to identify which text corresponds to which scan; keeping the same base name for the image and the recognized text solves the problem from the start.

#Insert it into a local pipeline

In a local document-processing setup, Tesseract is used in only one place: just before indexing, on documents without a text layer. The test to automate is simple—if text extraction from a PDF returns fewer than a few dozen characters per page (a threshold to adjust for your corpus), it is probably a scan and should be sent to OCR. Without this test, entire documents enter the index silently empty, and no one notices until a user is surprised.

Deskewing and binarization should be automated in this same pipeline rather than left to the judgment of each page: a script that detects the dominant angle and corrects it before calling the engine avoids many of the unpleasant surprises described above. On a homogeneous batch—the same scanner and the same document type every time—you configure this preprocessing once and it then runs unattended.

→
Measure before optimizing
Before tuning anything, assemble a small batch of twenty to thirty pages representative of the real corpus and record the error rate before and after each change. This is the only way to know whether deskewing, binarization, or a change in segmentation mode actually helped on your documents rather than on an isolated example.

#FAQ

Is Tesseract free?+
Yes, it is free software under the Apache 2.0 license, usable commercially without royalties, and works offline with no per-page fees or API key. The official repository has more than 70,000 stars on GitHub, indicating a large user base and a project that is still actively maintained in 2026, with a stable version released in July.
How do you make it read French?+
By installing the French language data file and explicitly requesting it during the call, with the -l fra option. Without it, the engine reads in English and mangles accents, turning an “é” into an approximate character. You can combine multiple languages at once, for example fra+eng for a bilingual document.
Why is my output unreadable?+
Most often, image quality: insufficient resolution, a skewed page, or low contrast. A document tilted by only three to five degrees can reduce the reading rate from 100% to about 31% in tests by the vendor Koncile. Straightening, binarizing, and targeting 300 dots per inch generally helps more than engine settings.
Can Tesseract read tables?+
It renders the text, not the structure, and the official documentation mentions a known issue with tables without custom segmentation. Rows and columns are lost, and an amount may end up associated with the wrong label. For tables, you need a layout analysis tool such as PaddleOCR or a vision model capable of reconstructing cells before trusting the extracted figures.
Should you choose a vision model?+
For complex layouts or handwriting, yes: FastOCR, a competing service, puts Tesseract v5 at between 25 and 40% accuracy on handwriting, a figure that should be treated with caution. On clean printed text, Tesseract remains faster and lighter, and it invents fewer plausible values: it confuses characters, whereas a model can produce a false but credible value.
What is page segmentation mode?+
A setting that tells the engine the document layout before reading it: full page, single column, isolated line, isolated word, or scattered text with no particular order. This is the first setting to change when a crop that visibly contains text returns nothing. Adding a thin white border around a crop that is too tightly trimmed also helps, according to the documentation.
How do you make a scanned PDF searchable with Tesseract?+
For an image, add pdf at the end of the command: Tesseract produces a PDF with a hidden text layer. For a multipage PDF, use OCRmyPDF, which relies on Tesseract, rasterizes the pages, and then adds this layer to the document. In both cases, specify the language with -l fra; otherwise English is assumed and accents are lost.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.