PaddleOCR: the OCR that understands page
PaddleOCR is an open-source OCR toolkit (Apache 2.0 license, more than 90,000 stars on GitHub) that detects text anywhere on a page and reconstructs tables. Since its PaddleOCR-VL variant, a vision model of approximately 0.9 billion parameters has reached 96.33% on the reference benchmark OmniDocBench v1.6—enough to replace line-by-line reading on real documents, at the cost of a heavier installation.
PaddleOCR does what a conventional optical recognition engine can't: detect text anywhere on a page, read dense handwriting, and reconstruct a table's structure. It's heavier than the legacy engine, and that's precisely what's needed for important documents — invoices, forms, reports, and multilingual scans.
#Detection first: what it changes
PaddleOCR is an open-source OCR toolkit under the Apache 2.0 license that does not read a page line by line: it first finds where the text is, then reads each region, and can then reconstruct the structure of the page and its tables. That makes it robust on invoices, forms, skewed scans, and multilingual documents, where an engine like Tesseract confidently produces nonsense. The project comes in two families: a modular pipeline (PP-OCRv6 for reading, PP-StructureV3 for structure) and PaddleOCR-VL, a vision model of approximately 0.9 billion parameters that converts a page to Markdown in a single pass and scores 96.33% on OmniDocBench v1.6 according to its vendor. The tradeoff is a heavier installation, with PaddlePaddle and model weights, plus documented GPU requirements for the VL variant. For clean text at scale, a lightweight engine remains simpler.
A traditional engine assumes a page consists of lines of text arranged like a book. Real documents do not: an invoice has boxes, a form has fields, a presentation has text over images, a scan may be skewed, and a technical drawing has labels at angles.
PaddleOCR separates the problem. A detection model finds text regions wherever they are and returns their positions; a recognition model reads each region. Slanted text, a margin caption, and a number in a cell become three regions among others. That's what lets it handle documents where a line-by-line reader makes mistakes without indicating them.
#The pipeline steps
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
| Step | What it produces | Why it matters |
|---|---|---|
| Text detection | Frames around every text area | Nothing is missed because of an unusual position |
| Orientation classification | The correct orientation of each zone | Skewed scans are no longer a special case |
| Recognition | The string for each frame | The actual reading step |
| Layout analysis | The type of each zone: title, paragraph, figure, table | Chunking can follow the structure instead of a character count |
| Table recognition | Rows, columns, cells | Numbers retain the label that gives them meaning |
Not every step is required. Reading a few labels only requires detection and recognition; ingesting financial reports for document search justifies the full pipeline. PP-OCRv6, the recognition-model generation released in 2026, alone covers 50 languages in a unified model (Chinese, English, Japanese, and 46 languages using the Latin alphabet), with no need to switch models between languages.
#PaddleOCR-VL: the OCR model becoming a vision model
Since October 2025, the project has published a second family under the same name: PaddleOCR-VL, a compact vision-language model that replaces the entire five-step pipeline with a single pass. Version 1.6, released in late May 2026, has about 0.9 billion parameters and combines a dynamic-resolution visual encoder with a small language model. On OmniDocBench v1.6, the reference benchmark for converting documents to Markdown or JSON, it achieves a score of 96.33%, a level the project itself describes as a new state of the art.
What stands out is not just the score: it is the size of the model achieving it. A guide published by InsiderLLM sums up the situation in one sentence: a 0.9-billion-parameter model that outperforms a 72-billion-parameter model and GPT-4o on document OCR. The same article quantifies the memory cost of the general-purpose competitor—Qwen2.5-VL-72B needs 48 GB of VRAM or more in Q4 quantization—while PaddleOCR-VL, converted to GGUF and quantized to Q4_K_M, occupies roughly one to one and a half gigabytes, including the language model weights and visual projector. This GGUF route is recent: support for PaddleOCR-VL was integrated into llama.cpp in February 2026 (version b8110), and the available GGUF files come from the community, not the PaddleOCR team.
The component that produces the OmniDocBench score is not alone: alongside PaddleOCR-VL, the project maintains PP-StructureV3, a pipeline dedicated to converting complex PDFs into Markdown or JSON with the precise coordinates of every table cell and text block. Both components pursue the same goal—a real document converted cleanly—by two different paths: a single model for PaddleOCR-VL, and a chain of specialized components for PP-StructureV3. Whichever engine you use, it natively supports multi-GPU and multi-process inference, which matters when you need to process a corpus of several tens of thousands of pages within a reasonable timeframe.
#PaddleOCR or Tesseract
| PaddleOCR | Tesseract | |
|---|---|---|
| Installation | Python stack and model weights | A small binary |
| Clean scan in one column | Excellent | Excellent, and faster |
| Text anywhere on the page | Excellent | Low |
| Tables | Reconstructed structure | Aplatis |
| Non-Latin writing | Very good | Depends heavily on the language package |
| Document tilted by a few degrees | Corrected by orientation classification | Can sharply reduce read speed |
| GPU | Optional, major gain | Not used |
A well-designed pipeline uses both: simple pages go to the lightweight engine, complex pages to the more expensive engine. Routing costs nothing and saves hours on a large corpus. The line about skew is not incidental: Koncile, an invoice-extraction editor, measured in its own tests that Tesseract's reading rate dropped from 100% on an upright invoice to 31% when the scan was tilted by just 3 to 5 degrees—the exact case that PaddleOCR's orientation-classification step is designed to handle.
- Tesseract: the lightweight engine, and when it’s enough
- Docling: converting structured documents
- Extract invoice data end to end
- Analyze an image with a local vision model
#Use cases: when to switch to PaddleOCR
- Billing and accounting
- An invoice scan is rarely perfectly straight; the table structure (quantity, unit price, VAT) must remain legible so the amounts retain their meaning instead of ending up as a line of isolated numbers.
- Multilingual administrative documents
- Forms, identity documents, and correspondence in multiple scripts: PaddleOCR's broad language coverage eliminates the need to set up a different engine for each country or alphabet.
- Long reports for a RAG system
- A report spanning several dozen pages with titles, subtitles, and tables is best converted with its structure intact: chunking that respects sections is better than chunking by character count.
- Scanned archives in bulk
- Carelessly digitized boxes of paper documents, random orientation: the orientation-classification step absorbs most of the disorder before reading even begins.
- High volume and clean text
- Conversely, a high-volume stream of already well-structured receipts or statements is often better served by a lighter engine — see the comparison with Tesseract above.
#Install and run your first extraction
Installation uses pip, with one peculiarity: since the 3.x series, the paddleocr package is not sufficient on its own. The documentation requires installing the selected inference engine first (PaddlePaddle by default), followed by the paddleocr package. Weights for the detection, recognition, and, where applicable, layout models are downloaded the first time each pipeline is used.
- 01Install the libraryFirst install PaddlePaddle by following the official installation page (the CPU or GPU variant, depending on the machine), then run python -m pip install paddleocr. The base package accepts Python 3.8 or later; the extras for document analysis (paddleocr[doc-parser]) require Python 3.9 or later.
- 02Run an initial extractionThe paddleocr ocr command takes an image as input (option -i) and writes the results to the folder specified by --save_path. To convert a page to Markdown with PaddleOCR-VL, use the paddleocr doc_parser command with the same -i and --save_path options.
- 03Enable the useful stepsThe use_doc_orientation_classify, use_doc_unwarping, and use_textline_orientation options enable straightening for a skewed page or line. On clean scans, setting them to False speeds up processing; the official documentation also recommends disabling unnecessary features when inference is too slow.
#What it consumes
The project doesn't publish universal throughput per page: it depends on document density, enabled steps, and hardware, and the documentation recommends disabling unnecessary features or choosing lighter models when inference is slow. Measure on about twenty of your pages before extrapolating to a corpus. For PaddleOCR-VL, the official documentation lists GPU requirements NVIDIA (PaddlePaddle: compute capability 7.0 or higher and CUDA 11.8 or higher; vLLM: 8.0 or higher and CUDA 12.6 or higher) and also provides a path for x64 CPUs. The weights are downloaded once, then everything runs locally: no API, no per-page cost.
#The deciding argument: no fabrication
PaddleOCR reads pixels. It can misread a character, and a digit substitution is not always noticeable. A general-purpose vision model asked to transcribe a document may produce a well-formed, plausible value that is not on the page—and nothing in the output indicates it. PaddleOCR-VL is trickier: because it generates text instead of reading it zone by zone, it belongs to the same risk category as general-purpose vision models, even though it is trained for transcription. No system is immune to errors involving ambiguous characters.
For document review, accounting, or anything that must be auditable, this difference should guide your choice: a dedicated OCR engine for the numbers you will rely on, and a general-purpose vision model when you want the document explained rather than transcribed. For high-stakes documents, using both and comparing them remains a legitimate third option.
In practice, verification does not require rereading everything. A sample of a few dozen documents per batch, manually compared with the engine's output, is enough to detect systematic drift—a poorly framed field, an incorrectly identified language, or a table that is regularly split incorrectly—before it contaminates thousands of automatically processed pages.
- Source: official PaddleOCR repository on GitHub
- Source: PaddleOCR-VL-1.6 model card
- Source: independent local PaddleOCR-VL comparison
#FAQ
Is PaddleOCR free?+
Do you need a GPU?+
PaddleOCR or Tesseract?+
Can it extract tables?+
What exactly is PaddleOCR-VL?+
Does it work offline?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.