MarkItDown: Getting Any File Into Markdown a Model Can Read
MarkItDown converts Office files, PDFs, HTML and more into markdown with one command. It is fast, tiny and deliberately shallow — which makes it the right first step for some corpora and the wrong one for others.
Key takeaways
- MarkItDown is a Microsoft open-source project under the MIT license, with more than 187,000 GitHub stars as of late September 2026 — one of the largest document-conversion tools in the ecosystem by that measure.
- It converts a wide range of formats into markdown — Word, Excel, PowerPoint, PDF, HTML, images, audio — through one command-line tool and one Python API.
- Its purpose is feeding text to language models, so it optimizes for clean markdown structure (headings, lists, tables) rather than visual fidelity, and the project explicitly compares itself to
textract. - It is light and fast by default, but a first-party plugin (
markitdown-ocr) and optional Azure integrations add LLM-vision OCR, cloud layout analysis, and even audio/video support when the built-in converters are not enough. - The right mental model: MarkItDown for office documents and web content, a heavier layout-aware parser for scanned or visually complex PDFs.
Why markdown, and why one tool
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- 30-day refund
A model reads text. Feed it raw HTML and most of the context window goes on markup; feed it a stripped text dump and you lose the headings and lists that told the model what the document's structure was. Markdown is the compromise: light enough to cost almost nothing in tokens, structured enough that a heading stays a heading. The project's own README makes the same point: it targets consumption by text-analysis tools, not high-fidelity conversion for human readers, and names textract as its closest comparison — a useful anchor if you have used that library before.
The second argument is uniformity. A real corpus is a spreadsheet, three Word files, a slide deck and two PDFs. Without a single converter, ingestion becomes a pile of format-specific code, each with its own edge cases and its own way of failing silently on a file nobody tested. MarkItDown is developed by Microsoft, MIT-licensed, and had more than 187,000 GitHub stars at the end of September 2026 — a scale of adoption that says less about conversion quality in the abstract than about how common this exact ingestion problem is across teams building anything RAG-shaped.
Installing it
- Install with the full feature set —
pip install 'markitdown[all]'pulls every optional dependency; installing from source works the same way against the repository checkout. - Or install only what you need —
pip install 'markitdown[pdf, docx, pptx]'installs dependencies for just those formats, keeping the environment smaller. - Run it from the command line —
markitdown path-to-file.pdf > document.md, or pipe a file in:cat path-to-file.pdf | markitdown. - Or call it from Python —
from markitdown import MarkItDown; md = MarkItDown(); result = md.convert("test.xlsx"); print(result.markdown). - Or run it in Docker —
docker build -t markitdown:latest .thendocker run --rm -i markitdown:latest < file.pdf > output.md, useful for a stateless conversion step in a larger pipeline.
What it handles, and how well
| Format | Result | Comment |
|---|---|---|
| Word documents | Very good | Headings, lists and simple tables survive |
| Spreadsheets | Good | Sheets become markdown tables; formulas become their values |
| Presentations | Good | Slide text and notes, without the visual arrangement; can use an LLM for image descriptions |
| HTML | Good | Useful for pages you already have on disk |
| Images | Metadata, plus optional LLM description | EXIF metadata is read directly; the built-in LLM-vision path (pptx and image files) needs an OpenAI-compatible client |
| Text-layer PDFs | Variable | Fine on simple layouts, weak on multi-column pages and complex tables |
| Scanned PDFs and images | Needs OCR | No text layer means nothing to extract by default — see Tesseract, or the markitdown-ocr plugin below |
The Python API accepts llm_client and llm_model arguments to generate image descriptions with an LLM instead of leaving images undescribed — currently limited to PowerPoint and image files. It also exposes max_retries for the underlying OpenAI-compatible client, with a documented default of 2 retries; if every attempt fails, MarkItDown falls back to other applicable converters rather than aborting the whole file.
Extending it: plugins and Azure
Plugins are disabled by default; markitdown --list-plugins shows what is installed, and markitdown --use-plugins path-to-file.pdf turns them on. Third-party plugins are discoverable on GitHub under the #markitdown-plugin hashtag. The project's own markitdown-ocr plugin is the most relevant one for local RAG: it adds OCR to PDF, DOCX, PPTX and XLSX converters by extracting text from embedded images through LLM vision, reusing the same llm_client / llm_model pattern MarkItDown already uses for image descriptions, with no extra ML libraries or binary dependencies. Installing it is pip install markitdown-ocr plus an OpenAI-compatible client; without an llm_client, the plugin still loads but silently skips OCR and falls back to the standard converter.
Beyond plugins, two optional Azure integrations push quality further at the cost of a network call: Azure Document Intelligence (install extra [az-doc-intel], pass a docintel_endpoint) for cloud-based layout extraction, and Azure Content Understanding (extra [az-content-understanding]) for structured field extraction serialized as YAML front matter, plus the only supported path for video and higher-quality audio. Both are billable cloud calls, not local compute, and neither is required for the everyday case of clean office documents — they exist for the harder tail of a corpus, not as a default you should reach for first.
An official MCP server exists too. The markitdown-mcp package is explicitly in scope for the project alongside the CLI, which means an MCP-compatible agent — Claude Code, Cursor, or any other client — can call MarkItDown's conversion directly instead of shelling out to the CLI from custom glue code.
MarkItDown or a layout-aware parser
This is the only comparison that matters, and it is not about quality in the abstract but about your files.
| MarkItDown | Docling | |
|---|---|---|
| Weight | A small package, no model weights for common formats | Model-based, heavier, GPU helps |
| Speed | Very fast | Seconds per complex page |
| Office files | Excellent | Also good |
| Complex PDF layout | Weak by default | Its reason to exist |
| Table structure | Basic | Recovered properly |
| Scanned PDFs | Needs the OCR plugin or an external step | Handled natively |
| Right when | Born-digital office and web documents | Scanned or visually complex PDFs |
A sensible pipeline uses both: route by file type and by whether a PDF has a text layer. Cheap conversion for the 80% that is straightforward, expensive conversion for the rest. Choosing one tool for everything means either wasting compute on simple files or producing garbage on hard ones.
Worth knowing about MarkItDown's own project boundaries: its maintainers explicitly keep the repository to a Python library, a CLI and an MCP package, and list web servers, REST or HTTP APIs, and hosted conversion services as out of scope for contributions. In practice that means if you need a conversion microservice behind an API, you build it yourself on top of the library rather than expecting one from the project — a design choice that keeps the core small and matches its role as a component, not a product.
Placing it in a local stack
- Detect the type. Office or HTML goes to MarkItDown; a PDF gets checked for a text layer.
- Convert and cache. Conversion happens once per document version, never per query.
- Chunk along the markdown structure — by heading, keeping tables intact — rather than by character count.
- Embed and store, then answer from retrieved passages with a local model. The full picture is in what RAG is.
Readers assembling the rest of a local RAG stack around this ingestion step can compare the pieces side by side in our local RAG toolkit.
Spreadsheets deserve a decision. A large sheet converted wholesale becomes a wall of table rows that dominates retrieval and answers nothing well. Either convert sheet by sheet with the headers repeated per chunk, or query the data as data rather than as text.
Verdict
MarkItDown is the sensible default for the boring majority of a corpus: office files, web pages, anything born digital. It costs almost nothing to run and gets structure right where structure is explicit in the file format, and its plugin and Azure options give it a documented growth path for images, scanned pages, audio and video without switching tools entirely for the whole pipeline. Keep a heavier parser for scanned and visually complex PDFs, route by file type, and cache everything — the combination is both cheaper and better than either tool alone.
Sources
- MarkItDown repository on GitHub — license, stars, releases
- MarkItDown README — install options, plugins, Azure integrations, Python and CLI usage
- markitdown-mcp package — the official MCP server
Frequently asked questions
Is MarkItDown free?
Yes. It is an MIT-licensed open-source project maintained by Microsoft, with more than 187,000 GitHub stars, usable commercially and locally with no API key for the built-in converters. The optional Azure Document Intelligence and Content Understanding integrations are billable cloud services, but they are entirely optional and never required for everyday office files.
Does it need a GPU?
No. The common formats are parsed structurally, not with vision models, so a CPU handles the default conversion path comfortably even on modest hardware. Optional LLM-based features — image descriptions and the markitdown-ocr plugin — call out to whatever OpenAI-compatible client you configure, local or remote, rather than running a model in-process.
Can it read scanned PDFs?
Not with the default converters — a scan has no text layer to extract, so there is nothing for the structural parser to find. The project now ships a first-party markitdown-ocr plugin that adds OCR via LLM vision to PDF, DOCX, PPTX and XLSX files, or you can run Tesseract yourself first and convert the recognized text afterward.
MarkItDown or Docling?
MarkItDown for office documents and web content, where it is smaller and faster with no model weights required for the common formats. Docling for scanned or visually complex PDFs, where layout and table recovery decide whether the output is usable at all. Many local pipelines route by file type and use both rather than picking one tool for the entire corpus.
Does it preserve tables?
Simple tables become markdown tables cleanly, and spreadsheet formulas are converted to their computed values rather than left as formula text. Merged cells and complex layouts degrade in the default converters, which is exactly where a layout-aware parser like Docling earns its extra weight and slower conversion time per page.
Is the output good enough for RAG?
For born-digital office and web files, yes — and chunking along its markdown headings works noticeably better than splitting by raw character count, since a heading stays attached to the content it introduces. For messy PDFs, spot-check a few conversions by hand before trusting the whole corpus, or route those files to a heavier parser instead.
Can an AI agent call MarkItDown directly?
Yes, through the official markitdown-mcp package, which the project lists as in scope for the repository alongside its command-line interface. An MCP-compatible agent such as Claude Code or Cursor can request a file conversion the same way it would call any other tool, instead of shelling out to the CLI through custom integration code written by hand.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.