Intermediate 12 minVision

PaddleOCR: the OCR that understands page

Direct response

PaddleOCR is an open-source OCR toolkit (Apache 2.0 license, more than 90,000 stars on GitHub) that detects text anywhere on a page and reconstructs tables. Since its PaddleOCR-VL variant, a vision model of approximately 0.9 billion parameters has reached 96.33% on the reference benchmark OmniDocBench v1.6—enough to replace line-by-line reading on real documents, at the cost of a heavier installation.

PaddleOCR does what a conventional optical recognition engine can't: detect text anywhere on a page, read dense handwriting, and reconstruct a table's structure. It's heavier than the legacy engine, and that's precisely what's needed for important documents — invoices, forms, reports, and multilingual scans.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Detection first: what it changes

PaddleOCR is an open-source OCR toolkit under the Apache 2.0 license that does not read a page line by line: it first finds where the text is, then reads each region, and can then reconstruct the structure of the page and its tables. That makes it robust on invoices, forms, skewed scans, and multilingual documents, where an engine like Tesseract confidently produces nonsense. The project comes in two families: a modular pipeline (PP-OCRv6 for reading, PP-StructureV3 for structure) and PaddleOCR-VL, a vision model of approximately 0.9 billion parameters that converts a page to Markdown in a single pass and scores 96.33% on OmniDocBench v1.6 according to its vendor. The tradeoff is a heavier installation, with PaddlePaddle and model weights, plus documented GPU requirements for the VL variant. For clean text at scale, a lightweight engine remains simpler.

A traditional engine assumes a page consists of lines of text arranged like a book. Real documents do not: an invoice has boxes, a form has fields, a presentation has text over images, a scan may be skewed, and a technical drawing has labels at angles.

PaddleOCR separates the problem. A detection model finds text regions wherever they are and returns their positions; a recognition model reads each region. Slanted text, a margin caption, and a number in a cell become three regions among others. That's what lets it handle documents where a line-by-line reader makes mistakes without indicating them.

#The pipeline steps

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
What each step produces
StepWhat it producesWhy it matters
Text detectionFrames around every text areaNothing is missed because of an unusual position
Orientation classificationThe correct orientation of each zoneSkewed scans are no longer a special case
RecognitionThe string for each frameThe actual reading step
Layout analysisThe type of each zone: title, paragraph, figure, tableChunking can follow the structure instead of a character count
Table recognitionRows, columns, cellsNumbers retain the label that gives them meaning

Not every step is required. Reading a few labels only requires detection and recognition; ingesting financial reports for document search justifies the full pipeline. PP-OCRv6, the recognition-model generation released in 2026, alone covers 50 languages in a unified model (Chinese, English, Japanese, and 46 languages using the Latin alphabet), with no need to switch models between languages.

#PaddleOCR-VL: the OCR model becoming a vision model

Since October 2025, the project has published a second family under the same name: PaddleOCR-VL, a compact vision-language model that replaces the entire five-step pipeline with a single pass. Version 1.6, released in late May 2026, has about 0.9 billion parameters and combines a dynamic-resolution visual encoder with a small language model. On OmniDocBench v1.6, the reference benchmark for converting documents to Markdown or JSON, it achieves a score of 96.33%, a level the project itself describes as a new state of the art.

What stands out is not just the score: it is the size of the model achieving it. A guide published by InsiderLLM sums up the situation in one sentence: a 0.9-billion-parameter model that outperforms a 72-billion-parameter model and GPT-4o on document OCR. The same article quantifies the memory cost of the general-purpose competitor—Qwen2.5-VL-72B needs 48 GB of VRAM or more in Q4 quantization—while PaddleOCR-VL, converted to GGUF and quantized to Q4_K_M, occupies roughly one to one and a half gigabytes, including the language model weights and visual projector. This GGUF route is recent: support for PaddleOCR-VL was integrated into llama.cpp in February 2026 (version b8110), and the available GGUF files come from the community, not the PaddleOCR team.

The component that produces the OmniDocBench score is not alone: alongside PaddleOCR-VL, the project maintains PP-StructureV3, a pipeline dedicated to converting complex PDFs into Markdown or JSON with the precise coordinates of every table cell and text block. Both components pursue the same goal—a real document converted cleanly—by two different paths: a single model for PaddleOCR-VL, and a chain of specialized components for PP-StructureV3. Whichever engine you use, it natively supports multi-GPU and multi-process inference, which matters when you need to process a corpus of several tens of thousands of pages within a reasonable timeframe.

i
Two products, one name
PaddleOCR (the modular five-step pipeline) and PaddleOCR-VL (the single-pass vision-language model) are two separate tools published by the same team. The former suits high-volume workloads where you want to control every step; the latter suits complex documents that you want to convert directly into structured Markdown without building a pipeline.

#PaddleOCR or Tesseract

Two budgets, not two rivals (qualitative assessments)
PaddleOCRTesseract
InstallationPython stack and model weightsA small binary
Clean scan in one columnExcellentExcellent, and faster
Text anywhere on the pageExcellentLow
TablesReconstructed structureAplatis
Non-Latin writingVery goodDepends heavily on the language package
Document tilted by a few degreesCorrected by orientation classificationCan sharply reduce read speed
GPUOptional, major gainNot used

A well-designed pipeline uses both: simple pages go to the lightweight engine, complex pages to the more expensive engine. Routing costs nothing and saves hours on a large corpus. The line about skew is not incidental: Koncile, an invoice-extraction editor, measured in its own tests that Tesseract's reading rate dropped from 100% on an upright invoice to 31% when the scan was tilted by just 3 to 5 degrees—the exact case that PaddleOCR's orientation-classification step is designed to handle.

#Use cases: when to switch to PaddleOCR

Billing and accounting
An invoice scan is rarely perfectly straight; the table structure (quantity, unit price, VAT) must remain legible so the amounts retain their meaning instead of ending up as a line of isolated numbers.
Multilingual administrative documents
Forms, identity documents, and correspondence in multiple scripts: PaddleOCR's broad language coverage eliminates the need to set up a different engine for each country or alphabet.
Long reports for a RAG system
A report spanning several dozen pages with titles, subtitles, and tables is best converted with its structure intact: chunking that respects sections is better than chunking by character count.
Scanned archives in bulk
Carelessly digitized boxes of paper documents, random orientation: the orientation-classification step absorbs most of the disorder before reading even begins.
High volume and clean text
Conversely, a high-volume stream of already well-structured receipts or statements is often better served by a lighter engine — see the comparison with Tesseract above.

#Install and run your first extraction

Installation uses pip, with one peculiarity: since the 3.x series, the paddleocr package is not sufficient on its own. The documentation requires installing the selected inference engine first (PaddlePaddle by default), followed by the paddleocr package. Weights for the detection, recognition, and, where applicable, layout models are downloaded the first time each pipeline is used.

  1. 01
    Install the library
    First install PaddlePaddle by following the official installation page (the CPU or GPU variant, depending on the machine), then run python -m pip install paddleocr. The base package accepts Python 3.8 or later; the extras for document analysis (paddleocr[doc-parser]) require Python 3.9 or later.
  2. 02
    Run an initial extraction
    The paddleocr ocr command takes an image as input (option -i) and writes the results to the folder specified by --save_path. To convert a page to Markdown with PaddleOCR-VL, use the paddleocr doc_parser command with the same -i and --save_path options.
  3. 03
    Enable the useful steps
    The use_doc_orientation_classify, use_doc_unwarping, and use_textline_orientation options enable straightening for a skewed page or line. On clean scans, setting them to False speeds up processing; the official documentation also recommends disabling unnecessary features when inference is too slow.
Terminal (after installing PaddlePaddle)
python -m pip install paddleocr
paddleocr ocr -i ./facture.jpg --use_doc_orientation_classify True --use_textline_orientation True --save_path ./output
→
Immediately usable output
The result includes the recognized text, the coordinates of each region, and, when structure detection is enabled, a Markdown or JSON export ready to be indexed for document search or sent to a language model. There's no need to write a custom parser to determine which cell belongs to which table.

#What it consumes

The project doesn't publish universal throughput per page: it depends on document density, enabled steps, and hardware, and the documentation recommends disabling unnecessary features or choosing lighter models when inference is slow. Measure on about twenty of your pages before extrapolating to a corpus. For PaddleOCR-VL, the official documentation lists GPU requirements NVIDIA (PaddlePaddle: compute capability 7.0 or higher and CUDA 11.8 or higher; vLLM: 8.0 or higher and CUDA 12.6 or higher) and also provides a path for x64 CPUs. The weights are downloaded once, then everything runs locally: no API, no per-page cost.

!
The GPU is shared
On a machine that also serves a language model, an OCR campaign and inference compete for the same memory. Start batch ingestion when no one is querying the system, and cache the result: a document is converted once, not for every question.

#The deciding argument: no fabrication

PaddleOCR reads pixels. It can misread a character, and a digit substitution is not always noticeable. A general-purpose vision model asked to transcribe a document may produce a well-formed, plausible value that is not on the page—and nothing in the output indicates it. PaddleOCR-VL is trickier: because it generates text instead of reading it zone by zone, it belongs to the same risk category as general-purpose vision models, even though it is trained for transcription. No system is immune to errors involving ambiguous characters.

For document review, accounting, or anything that must be auditable, this difference should guide your choice: a dedicated OCR engine for the numbers you will rely on, and a general-purpose vision model when you want the document explained rather than transcribed. For high-stakes documents, using both and comparing them remains a legitimate third option.

In practice, verification does not require rereading everything. A sample of a few dozen documents per batch, manually compared with the engine's output, is enough to detect systematic drift—a poorly framed field, an incorrectly identified language, or a table that is regularly split incorrectly—before it contaminates thousands of automatically processed pages.

#FAQ

Is PaddleOCR free?+
Yes. The project is released under the Apache 2.0 license, completely free including for commercial use, and has more than 90,000 stars on GitHub. It runs locally, with no API key or per-page cost. Still, check the license of the specific models you download, as some research variants may have different terms.
Do you need a GPU?+
Not to get started: the standard pipeline runs on the processor, and the PaddleOCR-VL documentation provides a path for an x64 processor. A NVIDIA GPU remains the best-documented option and becomes useful as soon as the volume exceeds a few thousand pages. The project does not publish CPU time per page: measure it on your documents.
PaddleOCR or Tesseract?+
Tesseract for clean, single-column text at high volume: it is faster and lighter as long as the scan remains straight. PaddleOCR for messy layouts, skewed scans, forms, tables, and non-Latin writing, where the lightweight engine quickly loses reliability.
Can it extract tables?+
Yes, structure recognition is part of the toolkit, through the standard pipeline or PP-StructureV3: rows, columns, and cells are reconstructed instead of being flattened into a line of numbers. No benchmark score guarantees the result on your own tables: manually check a few documents, especially those with merged cells, before trusting the entire pipeline.
What exactly is PaddleOCR-VL?+
A vision-language model with around 0.9 billion parameters, published by the same team as the classic OCR pipeline. It converts a page directly into Markdown or structured JSON in a single pass, with a score of 96.33% on the OmniDocBench v1.6 benchmark, without building a separate detection-recognition-structure pipeline.
Does it work offline?+
Yes, once the weights have been downloaded. Nothing leaves the machine afterward, which is the main reason to choose it for confidential documents: invoices, medical records, or legal documents remain on your disk from the first byte read to the last extracted character.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.