Intermediate 11 minStack

Docling: convert PDFs for AI locale

Direct response

Docling is an open-source library (MIT license, an IBM Research project hosted by the LF AI & Data Foundation) that converts PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, and audio into structured markdown or JSON, reconstructing layout, reading order, and tables. It runs entirely locally, with or without a GPU, and connects directly to LangChain, LlamaIndex, Crew AI, or Haystack to feed a document RAG pipeline.

PDF is the hardest input in any local document pipeline: two columns read across, a header that splits a sentence in two, a table of figures reduced to a column of numbers without their labels. Docling is an open-source library published by IBM Research, now hosted by the LF AI & Data Foundation, that analyzes the layout, reconstructs reading order, and retrieves table structure, then exports everything to markdown, HTML, or JSON. All on your machine, with no API — precisely the constraint when the documents are contracts or medical records.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#The step that determines everything else

When a document assistant gives a poor answer, we blame the model. The cause is almost always upstream. A report exported from a corporate template and passed through a naive text extractor becomes a stream where the page header interrupts a paragraph, a two-column layout is read horizontally, and a results table turns into a sequence of orphaned numbers. These fragments are then embedded, retrieved, and presented to the model as facts: it reads nonsense and confidently repeats it. No embedding model, reranker, or prompt can fix a failed conversion upstream in the pipeline—this is what most local RAG tutorials leave unsaid as they focus on choosing the language model.

Docling tackles this overlooked step directly. IBM Research released the project as open source in July 2024 under the MIT license, allowing commercial use without royalties or copyleft. It quickly surpassed 10,000 stars on GitHub and ranked among the world’s most-followed repositories in early 2025; by late September 2026, the official repository had more than 68,000 stars, and its latest stable version, v2.130.0, had been released on September 22, 2026.

#What Docling does

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Layout analysis
Identifies the areas of a page and their type, so a caption does not stick to a paragraph and a footer does not appear in the middle of a sentence.
Reading order
Reconstructs the sequence a human would follow, finally making documents arranged in columns usable.
Table structure
It finds rows, columns, and merged cells and exports them as real tables rather than lines of text.
Optical recognition when needed
Scanned pages without a text layer go through OCR instead of being ingested empty.
Chart understanding
Pie charts, bar charts, and line graphs can be converted into data tables or textual descriptions rather than simply ignored.
A single document model
One common internal representation, several export formats (Markdown, HTML, DocTags, lossless JSON): the rest of the pipeline does not care whether the input was a PDF or an office file.

Beyond PDF, Docling supports DOCX, PPTX, XLSX, HTML, EPUB, images (PNG, TIFF, JPEG…), emails (EML, MSG), and even audio (WAV, MP3) through a transcription pipeline—which matters because a real corpus is never homogeneous. For integrations, the library connects in a few lines to LangChain, LlamaIndex, Crew AI, and Haystack, avoiding the need to write the conversion code yourself in an agentic pipeline.

#The tables, the real reason to bother

In a professional document, numbers almost always live in tables, and that’s where naive extraction fails most badly. A row that becomes “Paris 12 480 3.2” has lost the column headings that gave it meaning. Search then returns a seemingly correct snippet, the model invents the relationship between the values, and the answer is wrong in a way that is difficult to detect.

Docling delegates this task to a dedicated model, TableFormer, which encodes table structure using a specialized vocabulary called OTSL (Optimized Table Structure Language) and correctly handles merged cells and multi-level headers. According to the research paper behind this format, OTSL expresses in a handful of tokens what an equivalent HTML representation expresses with more than 28 tokens, shortening the average sequence to predict by about half and cutting inference time in half compared with a model that generates HTML—an architectural detail invisible to the end user, but one that explains why Docling’s table recognition remains usable on large volumes without a GPU dedicated to this task alone. Two modes are available in the pipeline options: FAST, which is faster but less accurate on complex tables, and ACCURATE, recommended whenever tables contain merged cells or multiple header levels. Preserving the structure, even in markdown, maintains the link between each value and its label. For a corpus where the expected answers are numbers, that alone justifies using a heavier converter than a simple text extractor.

!
Manually check five documents
Before ingesting an entire corpus, convert five representative documents and read the resulting markdown, paying close attention to tables and page breaks. Ten minutes here can save a week of blaming the model.

#Available OCR engines

Docling does not bundle a single OCR engine: it orchestrates several interchangeable engines depending on the document type. EasyOCR and Tesseract (via tesserocr or the command line) cover most cases; RapidOCR accepts custom models; OcrMac uses native macOS recognition when available. The choice is made in the pipeline options, not in your application's code, so you can switch engines without changing the ingestion logic.

In practice, a production ingestion pipeline almost always starts by testing Tesseract on a sample: if it produces clean text, there’s no reason to pay EasyOCR’s cost. Switching to EasyOCR is mainly justified for degraded scans, partially handwritten forms, or languages where Tesseract falls short. RapidOCR and OcrMac remain niche choices, reserved respectively for an already-trained in-house OCR model and an isolated Mac without any external dependencies to install.

Choose an OCR engine based on the document
EngineStrengthTypical use case
TesseractFast on clean textDigital documents already scanned cleanly
EasyOCRMore robust on degraded scans and non-standard writing, with optional GPU support via use_gpuArchives, forms, heterogeneous corpora
RapidOCRAccepts custom modelsSpecific needs (rare language, specialized domain)
OcrMacUses macOS's native engine, with no additional dependenciesMac workstation, small volumes

#Chunking that follows the structure

Here is something most local RAG guides do not explain: Docling does not merely convert documents; it also offers chunking suited to what it has just reconstructed. Its HybridChunker starts from the document hierarchy (titles, sections), then adjusts each chunk’s size to the actual tokenizer of the selected embedding model: overly long blocks are split at element boundaries rather than in the middle of a sentence, while short blocks sharing the same title are merged. The provided tokenizer must be aligned with that of the downstream embedding model; otherwise, the actual chunk size in tokens no longer matches what the vector index expects. This is structure-based chunking, compared with character-block chunking, which ignores where sentences fall: see our guide to chunking strategies for the tradeoffs between the two approaches.

#The computational cost

Approximate scale by configuration
ConfigurationThroughputWhen it’s enough
CPU only, without OCRThe slowest: a few seconds per complex pageSmall corpora, occasional conversions
CPU-only, with OCREven slower, OCR dominatesA few scanned documents
With GPUSignificantly faster for page analysis, tables, and EasyOCR OCRThousands of pages, repeated ingestion

The practical consequence: convert in batches once, then keep the result. Reindexing only makes sense if the source changes. And on a machine that also serves a language model, both compete for the same GPU through the pipeline’s acceleration options: ingesting a corpus while users ask questions slows both down.

#Its place in the pipeline

  1. 01
    Convert
    Docling converts your files to Markdown or structured JSON, with tables intact.
  2. 02
    Split
    With the HybridChunker, following the retrieved structure and the embedding model's tokenizer instead of splitting every thousand characters. This is where the converter pays for itself a second time.
  3. 03
    Encode and store
    A local embedding model turns chunks into vectors, which are stored in a vector database such as Qdrant.
  4. 04
    Answer
    A local model writes from the retrieved passages, through Ollama or a local inference server.

#Where it still gets stuck

Poor-quality scans
OCR accuracy on a skewed photocopy is a physical limitation, not a software one.
Highly graphical layouts
Magazines, text wrapped around a figure, forms: difficult for every converter.
Handwriting
Outside the scope of this tool category.
Complex graphics
Chart understanding covers common cases (pie charts, bar charts, line charts); a highly specific chart—map, technical diagram, or architecture diagram—may still be reproduced only through its caption, without the underlying values.
The upfront cost
Installing the layout, table, and OCR models requires several gigabytes of downloads on first launch; on an isolated machine without network access, you need to retrieve them beforehand.

#FAQ

Is Docling free?+
Yes: it is an open-source project published by IBM Research in July 2024 under the MIT license and now hosted by the LF AI & Data Foundation. The license allows commercial use without royalties or any obligation to republish your own code, and the tool runs locally, with no API key or per-page or per-document conversion charges.
Do you need a GPU?+
No, Docling runs on the processor. However, its layout analysis, table recognition, and EasyOCR engine rely on models that can be enabled on the GPU through accelerator_options: a GPU greatly reduces conversion time for a large corpus. For a few dozen documents, the processor is more than sufficient.
Can it read scanned PDFs?+
Yes, through one of its OCR engines (EasyOCR, Tesseract, RapidOCR, or OcrMac, depending on the platform), which must be enabled for documents without a text layer. Quality then depends on the scan: a sharp document at 300 dots per inch works well, while a skewed photocopy performs much worse.
What is the difference between TableFormer's FAST and ACCURATE modes?+
FAST prioritizes speed and works well for simple tables with a single header row. ACCURATE is slower but recommended whenever the document contains merged cells or multi-level headers, which is common in financial reports or technical data sheets.
Docling or a simple text extractor?+
A simple extractor is faster and works well for clean, single-column text. Docling is justified for complex layouts, tables, and heterogeneous corpora that mix PDFs, Office files, and scans—that is, most real-world enterprise documents.
Does Docling integrate with LangChain or LlamaIndex?+
Yes, the project provides ready-to-use integrations for LangChain, LlamaIndex, Crew AI, and Haystack. In practice, this saves you from writing the connector between document conversion and the rest of the RAG pipeline yourself: the document converted by Docling arrives directly in the format expected by these frameworks' document loaders, ready for splitting and indexing.
Does data leave my machine?+
No, the conversion is entirely local once the layout, table recognition, and OCR models have been downloaded: no network call is required while processing a document. That is precisely the point for confidential documents—contracts, medical records, or accounting documents—that must remain on your computer or on a server you control, without passing through a third-party API.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.