Docling: convert PDFs for AI locale
Docling is an open-source library (MIT license, an IBM Research project hosted by the LF AI & Data Foundation) that converts PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, and audio into structured markdown or JSON, reconstructing layout, reading order, and tables. It runs entirely locally, with or without a GPU, and connects directly to LangChain, LlamaIndex, Crew AI, or Haystack to feed a document RAG pipeline.
PDF is the hardest input in any local document pipeline: two columns read across, a header that splits a sentence in two, a table of figures reduced to a column of numbers without their labels. Docling is an open-source library published by IBM Research, now hosted by the LF AI & Data Foundation, that analyzes the layout, reconstructs reading order, and retrieves table structure, then exports everything to markdown, HTML, or JSON. All on your machine, with no API — precisely the constraint when the documents are contracts or medical records.
#The step that determines everything else
When a document assistant gives a poor answer, we blame the model. The cause is almost always upstream. A report exported from a corporate template and passed through a naive text extractor becomes a stream where the page header interrupts a paragraph, a two-column layout is read horizontally, and a results table turns into a sequence of orphaned numbers. These fragments are then embedded, retrieved, and presented to the model as facts: it reads nonsense and confidently repeats it. No embedding model, reranker, or prompt can fix a failed conversion upstream in the pipeline—this is what most local RAG tutorials leave unsaid as they focus on choosing the language model.
Docling tackles this overlooked step directly. IBM Research released the project as open source in July 2024 under the MIT license, allowing commercial use without royalties or copyleft. It quickly surpassed 10,000 stars on GitHub and ranked among the world’s most-followed repositories in early 2025; by late September 2026, the official repository had more than 68,000 stars, and its latest stable version, v2.130.0, had been released on September 22, 2026.
#What Docling does
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- Layout analysis
- Identifies the areas of a page and their type, so a caption does not stick to a paragraph and a footer does not appear in the middle of a sentence.
- Reading order
- Reconstructs the sequence a human would follow, finally making documents arranged in columns usable.
- Table structure
- It finds rows, columns, and merged cells and exports them as real tables rather than lines of text.
- Optical recognition when needed
- Scanned pages without a text layer go through OCR instead of being ingested empty.
- Chart understanding
- Pie charts, bar charts, and line graphs can be converted into data tables or textual descriptions rather than simply ignored.
- A single document model
- One common internal representation, several export formats (Markdown, HTML, DocTags, lossless JSON): the rest of the pipeline does not care whether the input was a PDF or an office file.
Beyond PDF, Docling supports DOCX, PPTX, XLSX, HTML, EPUB, images (PNG, TIFF, JPEG…), emails (EML, MSG), and even audio (WAV, MP3) through a transcription pipeline—which matters because a real corpus is never homogeneous. For integrations, the library connects in a few lines to LangChain, LlamaIndex, Crew AI, and Haystack, avoiding the need to write the conversion code yourself in an agentic pipeline.
#The tables, the real reason to bother
In a professional document, numbers almost always live in tables, and that’s where naive extraction fails most badly. A row that becomes “Paris 12 480 3.2” has lost the column headings that gave it meaning. Search then returns a seemingly correct snippet, the model invents the relationship between the values, and the answer is wrong in a way that is difficult to detect.
Docling delegates this task to a dedicated model, TableFormer, which encodes table structure using a specialized vocabulary called OTSL (Optimized Table Structure Language) and correctly handles merged cells and multi-level headers. According to the research paper behind this format, OTSL expresses in a handful of tokens what an equivalent HTML representation expresses with more than 28 tokens, shortening the average sequence to predict by about half and cutting inference time in half compared with a model that generates HTML—an architectural detail invisible to the end user, but one that explains why Docling’s table recognition remains usable on large volumes without a GPU dedicated to this task alone. Two modes are available in the pipeline options: FAST, which is faster but less accurate on complex tables, and ACCURATE, recommended whenever tables contain merged cells or multiple header levels. Preserving the structure, even in markdown, maintains the link between each value and its label. For a corpus where the expected answers are numbers, that alone justifies using a heavier converter than a simple text extractor.
#Available OCR engines
Docling does not bundle a single OCR engine: it orchestrates several interchangeable engines depending on the document type. EasyOCR and Tesseract (via tesserocr or the command line) cover most cases; RapidOCR accepts custom models; OcrMac uses native macOS recognition when available. The choice is made in the pipeline options, not in your application's code, so you can switch engines without changing the ingestion logic.
In practice, a production ingestion pipeline almost always starts by testing Tesseract on a sample: if it produces clean text, there’s no reason to pay EasyOCR’s cost. Switching to EasyOCR is mainly justified for degraded scans, partially handwritten forms, or languages where Tesseract falls short. RapidOCR and OcrMac remain niche choices, reserved respectively for an already-trained in-house OCR model and an isolated Mac without any external dependencies to install.
| Engine | Strength | Typical use case |
|---|---|---|
| Tesseract | Fast on clean text | Digital documents already scanned cleanly |
| EasyOCR | More robust on degraded scans and non-standard writing, with optional GPU support via use_gpu | Archives, forms, heterogeneous corpora |
| RapidOCR | Accepts custom models | Specific needs (rare language, specialized domain) |
| OcrMac | Uses macOS's native engine, with no additional dependencies | Mac workstation, small volumes |
#Chunking that follows the structure
Here is something most local RAG guides do not explain: Docling does not merely convert documents; it also offers chunking suited to what it has just reconstructed. Its HybridChunker starts from the document hierarchy (titles, sections), then adjusts each chunk’s size to the actual tokenizer of the selected embedding model: overly long blocks are split at element boundaries rather than in the middle of a sentence, while short blocks sharing the same title are merged. The provided tokenizer must be aligned with that of the downstream embedding model; otherwise, the actual chunk size in tokens no longer matches what the vector index expects. This is structure-based chunking, compared with character-block chunking, which ignores where sentences fall: see our guide to chunking strategies for the tradeoffs between the two approaches.
#The computational cost
| Configuration | Throughput | When it’s enough |
|---|---|---|
| CPU only, without OCR | The slowest: a few seconds per complex page | Small corpora, occasional conversions |
| CPU-only, with OCR | Even slower, OCR dominates | A few scanned documents |
| With GPU | Significantly faster for page analysis, tables, and EasyOCR OCR | Thousands of pages, repeated ingestion |
The practical consequence: convert in batches once, then keep the result. Reindexing only makes sense if the source changes. And on a machine that also serves a language model, both compete for the same GPU through the pipeline’s acceleration options: ingesting a corpus while users ask questions slows both down.
#Its place in the pipeline
- 01ConvertDocling converts your files to Markdown or structured JSON, with tables intact.
- 02SplitWith the HybridChunker, following the retrieved structure and the embedding model's tokenizer instead of splitting every thousand characters. This is where the converter pays for itself a second time.
- 03Encode and storeA local embedding model turns chunks into vectors, which are stored in a vector database such as Qdrant.
- 04AnswerA local model writes from the retrieved passages, through Ollama or a local inference server.
- Store vectors in Qdrant
- Compare chunking strategies
- A turnkey application for chatting with your documents
- Special case: extracting data from invoices
- Tesseract alone: simpler OCR for clean text
- The QuelLLM local RAG kit: all components on one page
- Source: official Docling repository on GitHub
- Source: Docling pipeline options (OCR, TableFormer)
- Source: HybridChunker documentation
#Where it still gets stuck
- Poor-quality scans
- OCR accuracy on a skewed photocopy is a physical limitation, not a software one.
- Highly graphical layouts
- Magazines, text wrapped around a figure, forms: difficult for every converter.
- Handwriting
- Outside the scope of this tool category.
- Complex graphics
- Chart understanding covers common cases (pie charts, bar charts, line charts); a highly specific chart—map, technical diagram, or architecture diagram—may still be reproduced only through its caption, without the underlying values.
- The upfront cost
- Installing the layout, table, and OCR models requires several gigabytes of downloads on first launch; on an isolated machine without network access, you need to retrieve them beforehand.
#FAQ
Is Docling free?+
Do you need a GPU?+
Can it read scanned PDFs?+
What is the difference between TableFormer's FAST and ACCURATE modes?+
Docling or a simple text extractor?+
Does Docling integrate with LangChain or LlamaIndex?+
Does data leave my machine?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.