Tesseract OCR: Getting Text Out of Scans, Locally
A scanned PDF holds no text, only a picture of text. Tesseract is the free engine that fixes that, offline. Language data, the image quality that decides everything, and when a vision model would be worse.
Key takeaways
- A scanned PDF has no text layer. Indexed as-is it contributes nothing, and your document assistant will insist the file says nothing at all.
- Tesseract is the free, offline OCR engine for that step: Apache-licensed, available everywhere, tiny compared to any model-based alternative.
- Image quality decides the outcome far more than any setting. 300 dpi, deskewed, high contrast — get that right and the engine looks excellent; get it wrong and no flag saves you.
- Two options are worth knowing: the language data (it defaults to English and mangles everything else) and the page segmentation mode (the default assumes a full page and returns nothing on a single line of text).
- Its decisive advantage over a vision model: it does not hallucinate. Tesseract makes visible errors; a vision model can invent a perfectly formatted wrong number.
What it is, and what it is not
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- 30-day refund
Tesseract is an optical character recognition engine with a long history — decades of development, now open source under Apache 2.0. Since version 4 it recognises text lines with a recurrent neural network rather than character by character, which lifted its accuracy on ordinary documents substantially.
What it is not: a document understanding system. It returns text, optionally with positions. It will not tell you that a block is a heading, will not rebuild a table into rows and columns, and does not understand what it reads. For structured documents, OCR is one brick among several — the layout work belongs to a parser like Docling.
Language data: the first thing people get wrong
Tesseract reads English by default. Run it on a French, German or Spanish document without installing that language's data file and you get lost accents and near-miss words — from which many people conclude the engine is bad. Install the language pack and request it explicitly.
| Situation | What to do |
|---|---|
| Non-English document | Install that language's data and pass it at runtime |
| Mixed-language document | Pass several languages together; accuracy drops slightly, coverage improves |
| Numbers-only field | Restrict the allowed character set to digits |
| Unusual script | Check a language pack exists before planning the project around it |
Image quality does the heavy lifting
This is the lesson everyone learns late: on a mediocre scan no flag rescues the output, while ten lines of preprocessing turn unreadable text into clean text.
- Resolution — around 300 dpi is the target. Below 200 accuracy collapses; far above it you pay time for nothing.
- Deskew — a page tilted two degrees costs more accuracy than any engine tuning recovers.
- Contrast and binarisation — light grey text on a grey background should become black on white before it reaches the engine.
- Cropping — removing scanner borders and stray margins prevents phantom characters.
A PDF is not an image. Tesseract expects images, so a PDF must first be rasterised page by page — or handed to a wrapper that does both steps and writes the recognised text back into an invisible layer of the PDF. That is what "OCR my PDF" tools do, and it is usually what you want, because the file stays a PDF and becomes searchable.
Page segmentation: the setting that returns empty strings
Tesseract tries to work out the structure of what it is looking at before reading it. The default assumes a full page of text. Hand it a single line — a label, a receipt line, a narrow column — and it may return nothing at all, which reads like a failure and is a misconfiguration.
Telling it what the image contains (a block of text, a single line, a single word) is often the difference between an empty result and a perfect one. It is the first thing to change when a crop that clearly contains text produces nothing.
Tesseract or a vision model
| Criterion | Tesseract | Local vision model |
|---|---|---|
| Resources | CPU, tens of megabytes | GPU preferred, several gigabytes |
| Speed | Very fast | Much slower |
| Clean single-column text | Excellent | Equivalent, more expensive |
| Complex layout, tables | Weak | Much better |
| Understands content | No | Can answer questions about the page |
| Invents values | Never — it errs, it does not fabricate | Yes, plausibly formatted and wrong |
That last row decides most document-control use cases. Tesseract produces visible garbage when it fails; a vision model can produce a clean, confident, incorrect figure. On accounting or administrative records that is a difference in kind, not degree. For layout-heavy documents where you accept that risk, PaddleOCR and vision models are the stronger tools.
Where it belongs in a local stack
In a local document pipeline, Tesseract appears at exactly one point: immediately before indexing, on files with no text layer. The test to automate is simple — if extracting text from a PDF yields fewer than a few dozen characters per page, it is a scan and goes to OCR. Without that check, whole documents enter the index empty and nobody notices until a user asks why the assistant has never heard of a contract that is right there. The rest of the chain is the usual one, ending in a local model answering from retrieved passages — see what RAG is.
Verdict
Tesseract remains the sensible default for the OCR step: free, offline, fast, and predictable in how it fails. Spend your effort on the images rather than the flags, install the right language data, and tell it what shape the page is. Reach for a vision-based reader only when the layout genuinely defeats it — and never on documents where an invented number would go unnoticed.
Frequently asked questions
Is Tesseract free for commercial use?
Yes. It is open source under Apache 2.0, runs offline, and has no per-page cost.
Why is my OCR output garbage?
Image quality, in the overwhelming majority of cases: resolution below 200 dpi, a tilted page, or weak contrast. Deskew, binarise and target 300 dpi before touching any engine setting.
Can Tesseract read tables?
It returns text, not structure — rows and columns are lost. For tables you need a layout-aware parser or a vision model that recognises table structure.
How do I OCR a non-English document?
Install the data file for that language and request it explicitly at runtime. Without it the engine reads English and mangles accented characters.
Should I use a vision model instead?
For complex layouts, often yes. For clean text, Tesseract is faster, lighter and safer, because it does not fabricate plausible values — its mistakes are visible rather than convincing.
Does Tesseract work on PDFs directly?
Not directly; it takes images. Either rasterise the pages first, or use a wrapper that converts, recognises, and writes the text back into the PDF as an invisible layer.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.