BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-21

Tesseract OCR: Getting Text Out of Scans, Locally

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

A scanned PDF holds no text, only a picture of text. Tesseract is the free engine that fixes that, offline. Language data, the image quality that decides everything, and when a vision model would be worse.

By Mohamed Meguedmi·Last updated 2026-09-21·10 min read·Tested on Windows, macOS, Linux

Key takeaways

  • A scanned PDF has no text layer. Indexed as-is it contributes nothing, and your document assistant will insist the file says nothing at all.
  • Tesseract is the free, offline OCR engine for that step: Apache-licensed, available everywhere, tiny compared to any model-based alternative.
  • Image quality decides the outcome far more than any setting. 300 dpi, deskewed, high contrast — get that right and the engine looks excellent; get it wrong and no flag saves you.
  • Two options are worth knowing: the language data (it defaults to English and mangles everything else) and the page segmentation mode (the default assumes a full page and returns nothing on a single line of text).
  • Its decisive advantage over a vision model: it does not hallucinate. Tesseract makes visible errors; a vision model can invent a perfectly formatted wrong number.

What it is, and what it is not

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • 30-day refund

Tesseract is an optical character recognition engine with a long history — decades of development, now open source under Apache 2.0. Since version 4 it recognises text lines with a recurrent neural network rather than character by character, which lifted its accuracy on ordinary documents substantially.

What it is not: a document understanding system. It returns text, optionally with positions. It will not tell you that a block is a heading, will not rebuild a table into rows and columns, and does not understand what it reads. For structured documents, OCR is one brick among several — the layout work belongs to a parser like Docling.

Language data: the first thing people get wrong

Tesseract reads English by default. Run it on a French, German or Spanish document without installing that language's data file and you get lost accents and near-miss words — from which many people conclude the engine is bad. Install the language pack and request it explicitly.

SituationWhat to do
Non-English documentInstall that language's data and pass it at runtime
Mixed-language documentPass several languages together; accuracy drops slightly, coverage improves
Numbers-only fieldRestrict the allowed character set to digits
Unusual scriptCheck a language pack exists before planning the project around it

Image quality does the heavy lifting

This is the lesson everyone learns late: on a mediocre scan no flag rescues the output, while ten lines of preprocessing turn unreadable text into clean text.

  • Resolution — around 300 dpi is the target. Below 200 accuracy collapses; far above it you pay time for nothing.
  • Deskew — a page tilted two degrees costs more accuracy than any engine tuning recovers.
  • Contrast and binarisation — light grey text on a grey background should become black on white before it reaches the engine.
  • Cropping — removing scanner borders and stray margins prevents phantom characters.

A PDF is not an image. Tesseract expects images, so a PDF must first be rasterised page by page — or handed to a wrapper that does both steps and writes the recognised text back into an invisible layer of the PDF. That is what "OCR my PDF" tools do, and it is usually what you want, because the file stays a PDF and becomes searchable.

Page segmentation: the setting that returns empty strings

Tesseract tries to work out the structure of what it is looking at before reading it. The default assumes a full page of text. Hand it a single line — a label, a receipt line, a narrow column — and it may return nothing at all, which reads like a failure and is a misconfiguration.

Telling it what the image contains (a block of text, a single line, a single word) is often the difference between an empty result and a perfect one. It is the first thing to change when a crop that clearly contains text produces nothing.

Tesseract or a vision model

CriterionTesseractLocal vision model
ResourcesCPU, tens of megabytesGPU preferred, several gigabytes
SpeedVery fastMuch slower
Clean single-column textExcellentEquivalent, more expensive
Complex layout, tablesWeakMuch better
Understands contentNoCan answer questions about the page
Invents valuesNever — it errs, it does not fabricateYes, plausibly formatted and wrong

That last row decides most document-control use cases. Tesseract produces visible garbage when it fails; a vision model can produce a clean, confident, incorrect figure. On accounting or administrative records that is a difference in kind, not degree. For layout-heavy documents where you accept that risk, PaddleOCR and vision models are the stronger tools.

Where it belongs in a local stack

In a local document pipeline, Tesseract appears at exactly one point: immediately before indexing, on files with no text layer. The test to automate is simple — if extracting text from a PDF yields fewer than a few dozen characters per page, it is a scan and goes to OCR. Without that check, whole documents enter the index empty and nobody notices until a user asks why the assistant has never heard of a contract that is right there. The rest of the chain is the usual one, ending in a local model answering from retrieved passages — see what RAG is.

Verdict

Tesseract remains the sensible default for the OCR step: free, offline, fast, and predictable in how it fails. Spend your effort on the images rather than the flags, install the right language data, and tell it what shape the page is. Reach for a vision-based reader only when the layout genuinely defeats it — and never on documents where an invented number would go unnoticed.

Frequently asked questions

Is Tesseract free for commercial use?

Yes. It is open source under Apache 2.0, runs offline, and has no per-page cost.

Why is my OCR output garbage?

Image quality, in the overwhelming majority of cases: resolution below 200 dpi, a tilted page, or weak contrast. Deskew, binarise and target 300 dpi before touching any engine setting.

Can Tesseract read tables?

It returns text, not structure — rows and columns are lost. For tables you need a layout-aware parser or a vision model that recognises table structure.

How do I OCR a non-English document?

Install the data file for that language and request it explicitly at runtime. Without it the engine reads English and mangles accented characters.

Should I use a vision model instead?

For complex layouts, often yes. For clean text, Tesseract is faster, lighter and safer, because it does not fabricate plausible values — its mistakes are visible rather than convincing.

Does Tesseract work on PDFs directly?

Not directly; it takes images. Either rasterise the pages first, or use a wrapper that converts, recognises, and writes the text back into the PDF as an invisible layer.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.