Docling: Converting Real Documents Into Something a Model Can Read
PDFs are the hardest input in any local RAG pipeline. Docling parses layout, tables and reading order into structured markdown or JSON — locally, with no API, under an MIT license.
Key takeaways
- Docling is an MIT-licensed, open-source document conversion library originally released by IBM Research and now hosted under the LF AI & Data Foundation. Free for commercial use, no per-page fee.
- It does not dump text. It understands layout: reading order across multi-column pages, headings, and tables recovered as real tables instead of scrambled lines, via a dedicated model called TableFormer.
- It runs entirely on your machine and ships plug-and-play integrations for LangChain, LlamaIndex, Crew AI and Haystack, so it drops into an existing local RAG stack without custom glue code.
- Its HybridChunker splits documents along their own structure and then adjusts chunk size to the embedding model's tokenizer — a detail most ingestion tutorials skip entirely.
- Layout and table analysis are model-based, so they cost compute; a GPU is optional but turns hours into minutes on a large batch.
Why this step decides everything downstream
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- 30-day refund
Ask why a document chatbot answers badly and people blame the model. The cause usually sits upstream. A PDF exported from a report template becomes, under a naive text extractor, a stream where the page header interrupts a sentence, a two-column layout gets read across instead of down, and a table of figures turns into a column of numbers with no row labels attached. Those broken chunks are then embedded, retrieved and handed to the model as fact. The model does exactly what it was asked: it reads nonsense and states it with total confidence. No embedding model, no reranker and no prompt trick repairs a conversion step that already destroyed the structure.
Docling is IBM Research's answer to that specific problem. It was open-sourced in July 2024 under the MIT license, which permits commercial and personal use with no copyleft obligation. It crossed 10,000 GitHub stars within weeks of release and, according to the project's own GitHub statistics, the repository stood at more than 68,000 stars by late September 2026, with its latest stable release, v2.130.0, published on September 22, 2026 — a maintenance pace that matters when you are betting an ingestion pipeline on a library staying current.
What Docling actually parses
- Layout analysis — identifies the regions of a page and their type, so a caption is not glued to a paragraph and a footer does not appear mid-sentence.
- Reading order — reconstructs the sequence a human would follow, which is what makes multi-column filings and academic PDFs usable.
- Table structure recognition — recovers rows, columns and merged cells through TableFormer, and exports them as real tables rather than flattened text.
- Chart understanding — bar charts, pie charts and line plots can be converted into tables or code, with a generated description, instead of being silently dropped.
- OCR when needed — scanned pages with no text layer route through optical recognition rather than being ingested as empty.
Input coverage goes well beyond PDF: DOCX, PPTX, XLSX, HTML, EPUB, images (PNG, TIFF, JPEG), email formats (EML, MSG) and even audio (WAV, MP3) through a transcription pipeline. Output goes to Markdown, HTML, DocTags or lossless JSON. That range matters because a real corpus — a due-diligence data room, a claims folder, a compliance archive — is never a single file format.
Tables: the reason to bother with a heavier parser
Numbers in business documents almost always live inside tables, and that is where naive extraction fails hardest. A line that reads Chicago 12,480 3.2% has lost the column headers that gave it meaning. Retrieval then returns a plausible-looking chunk, the model invents the relationship between the values, and the answer is wrong in a way that is hard to catch in a demo.
Docling hands this job to TableFormer, which encodes table structure with a compact vocabulary called OTSL (Optimized Table Structure Language) instead of raw HTML tags. According to the research paper behind OTSL, the format needs only a handful of tokens for structures that would take an HTML representation 28 tokens or more, which shortens the average sequence length by about half and halves inference time compared with an HTML-based model — a low-level detail, but it is a large part of why table recognition stays usable on CPU-only batches. Two modes are exposed in the pipeline options: FAST, which trades some accuracy for speed on simple tables, and ACCURATE, recommended whenever a document has merged cells or multi-level headers — the norm in SEC filings, insurance schedules and pricing sheets.
Spot-check before trusting a corpus. Convert five representative documents by hand, read the markdown, and look specifically at the tables and at the page boundaries. Ten minutes here saves a week of blaming the model.
Choosing an OCR engine
Docling does not bundle a single OCR engine; it orchestrates several interchangeable ones through its pipeline options, not through application code. That separation means swapping engines is a config change, not a rewrite.
| Engine | Strength | Typical fit |
|---|---|---|
| Tesseract | Fast on clean, already-digital scans | Bulk conversion of tidy documents |
| EasyOCR | More robust on degraded scans; optional GPU via use_gpu | Archives, forms, mixed-quality corpora |
| RapidOCR | Accepts custom-trained models | Niche languages or domain-specific text |
| OcrMac | Uses macOS's native recognizer, no extra dependency | Small batches on a Mac workstation |
In practice, most production pipelines default to Tesseract on a sample first; if the output is clean, there is no reason to pay EasyOCR's extra compute. The switch to EasyOCR is usually triggered by skewed or low-contrast scans, not by document volume alone.
Chunking that respects the structure Docling just recovered
Something most local-RAG walkthroughs skip: Docling ships its own chunker, and it is not a fixed-character splitter. The HybridChunker starts from the document's own hierarchy — headings, sections — and then applies tokenization-aware refinements on top of that hierarchical chunking. Oversized chunks are split at item boundaries rather than mid-sentence, and undersized chunks that share the same heading are merged back together. Crucially, the tokenizer it uses should be aligned to the embedding model's own tokenizer; mismatch that and the chunk sizes reported in tokens no longer match what your vector index actually expects, which quietly degrades retrieval without throwing an error anywhere.
What it costs to run
| Setup | Throughput | When it is enough |
|---|---|---|
| CPU only, no OCR | Slowest; seconds per page on complex layouts | Small corpora, one-off conversions |
| CPU only, with OCR | Slower still — OCR dominates | A handful of scanned documents |
| GPU available | Much faster on layout, table and EasyOCR models via accelerator options | Thousands of pages, repeated ingestion |
The practical consequence: batch the work. Converting a corpus is a job you run once and cache, not something to redo on every query. On a machine that also serves the LLM, remember both want the same GPU — ingesting a large batch while users are asking questions slows down both, see how VRAM gets spent for why.
Where it sits in a RAG pipeline
- Convert — Docling turns files into structured markdown or JSON, tables intact.
- Chunk — the HybridChunker splits along that structure and the embedding model's tokenizer, not every N characters.
- Embed — a local embedding model turns chunks into vectors.
- Store and retrieve — a vector database, such as pgvector or FAISS, returns the closest passages.
- Answer — a local model writes the response from those passages, through Ollama or a local inference server.
For web pages rather than files, the equivalent first step is a crawler — see Firecrawl self-hosted. Readers assembling the rest of a local RAG stack around this choice can compare the pieces side by side in our local RAG toolkit.
Docling versus a plain text extractor or a cloud parsing API
| Option | Where it wins | Where it loses |
|---|---|---|
| Plain text extractor (pdftotext-style) | Fastest, zero setup, fine for clean single-column text | Destroys tables, columns and reading order on anything complex |
| Docling | Layout, table structure and OCR, fully local, MIT-licensed | Heavier install; model downloads needed before first run |
| Cloud document-parsing API | No local compute needed | Documents leave the machine — a blocker for regulated or confidential records |
The choice mostly comes down to one question: can the documents leave the machine? If the answer is no — contracts, medical records, anything under HIPAA or similar confidentiality rules — a cloud parser is off the table regardless of its accuracy, and Docling is one of the few options that keeps both layout understanding and table recovery entirely local.
Where it still struggles
- Poor scans. OCR accuracy on a skewed, faxed photocopy is a physical limit, not a software one.
- Heavily designed pages. Magazine layouts, text wrapped around figures, and dense forms remain hard for every parser, Docling included.
- Handwriting. Out of scope for this class of tool.
- Complex or unusual charts. Chart understanding covers the common cases (bar, pie, line); a highly specific diagram may still come back as caption only, with no underlying values.
- First-run setup. Layout, table and OCR models add up to several gigabytes on first launch — plan for that on an air-gapped machine.
Sources
Frequently asked questions
Is Docling free and open source?
Yes. IBM Research released it under the MIT license in July 2024, and the project is now hosted by the LF AI & Data Foundation. MIT permits commercial use with no copyleft obligation, and Docling runs locally with no API key and no per-document fee.
Does Docling need a GPU?
No, it runs on CPU. But layout analysis, table recognition and the EasyOCR engine are model-based, so a GPU cuts conversion time sharply on large batches through the pipeline's accelerator options. For a few dozen documents, CPU alone is fine and adds no real delay.
Which OCR engine should I pick?
Start with Tesseract on a sample; it is fast and works well on clean digital scans. Move to EasyOCR if scans are skewed, low-contrast or handwritten-adjacent — it costs more compute but is markedly more robust. RapidOCR and OcrMac stay niche picks for custom models or Mac-only workflows.
Does Docling replace a vector database?
No. Docling only converts and chunks documents; it does not store or search vectors itself. The chunks it produces still need an embedding model to turn them into vectors, and a vector database such as pgvector or FAISS to store and search them, before a local model can retrieve and answer from them in a RAG pipeline.
Docling or a simple PDF text extractor?
A simple extractor is faster and fine for clean, single-column text with no tables — think a plain contract with no exhibits. Docling earns its extra compute on multi-column layouts, tables and mixed corpora that combine PDFs, Office files and scans, in other words on most real business documents rather than on a textbook example.
Can documents stay entirely offline?
Yes, once the layout, table and OCR models are downloaded, conversion runs with no network call. That is the reason to pick Docling over a cloud parsing API for contracts, medical records or anything covered by confidentiality or compliance rules that forbid sending files to a third party.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.