DeepSeek-OCR locally: read its PDFs scanned
DeepSeek-OCR is a vision model with about 3 billion parameters, released under an MIT license and designed to turn an image of a page into structured text, including Markdown. This guide shows how to run deepseek ocr locally with Ollama: which runtime version to use, how much memory to plan for, how to split a scanned PDF into pages, and how to retrieve Markdown usable by a RAG. It ends with cases where a simpler tool does the job better.
#Why DeepSeek-OCR instead of traditional OCR
A scanned PDF contains no text, only images of text. A traditional OCR engine such as Tesseract extracts raw lines from it: it doesn't know that a block is a heading, it breaks tables, and it loses the reading order of a two-column page. DeepSeek-OCR approaches the problem differently. It's a vision-language model: it looks at the entire page, then directly generates Markdown with its headings, lists, tables, and formulas.
The model was released by DeepSeek in fall 2025 with a paper titled “Contexts Optical Compression.” The central idea is to represent a page with a small number of visual tokens, between 64 and 400 depending on the resolution mode, instead of the thousands of text tokens it would contain. The publisher claims decoding accuracy of approximately 97% when the compression ratio remains below ten. That is its figure, measured on its own datasets, not ours: the main takeaway is that the model is lightweight and fast for what it does.
For local use, three things matter. The model fits on an entry-level graphics card. It outputs Markdown that can be indexed as-is in a RAG pipeline. And it is available in the Ollama library, avoiding the need to install PyTorch, Flash Attention, and vLLM as required by the official repository. Full specifications are on the model page: https://quelllm.fr/modele/deepseek-ocr
#Prerequisites and memory requirements
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
The site’s model page provides VRAM guidelines by precision: approximately 2 GB in Q4, 4 GB in Q8, and 6 GB in FP16, for a context of 8,192 tokens. The actual download size depends on what Ollama packaged into the tag; the library’s tag page displays the file size, and that is the authoritative figure. Also account for context memory: a dense page converted to Markdown can produce several thousand output tokens.
- GPU NVIDIA
- Any card with 8 GB or more is comfortable. A RTX 3060 12 GB or a RTX 4070 12 GB can run the model in FP16 with room for context.
- Mac Apple Silicon
- Ollama uses unified memory. A Mac with 16 GB is enough for the model alone; plan on 24 GB if you want to keep a chat model loaded alongside it for RAG.
- Without a GPU
- Possible, but image encoding and Markdown generation run on the processor. For a few pages, it’s tolerable; for a batch of one hundred pages, expect minutes per page depending on the machine.
- Software
- Ollama up to date, poppler-utils for splitting PDFs, and Python 3 with the ollama library if you want to automate batch processing.
#Step 1: verify Ollama and retrieve the correct tag
DeepSeek-OCR arrived in the Ollama library in late 2025, after the model was released. It relies on a specialized vision encoder, DeepEncoder, which combines a local-window module with a global-attention module. Ollama had to add support for this architecture to its engine: an earlier runtime downloads the weights but refuses to load them, with a message saying that the model requires a newer version. The library page displays the minimum required version when applicable.
The show command is the only reliable source for knowing what you just downloaded: it reports the model family, parameter count, context length, and quantization level. If the latest tag and the 3b tag point to the same digest, it makes no difference which one you use; the guide uses 3b to remain explicit.
#Step 2: prepare the PDF pages
Ollama doesn't open PDFs. It accepts images, PNG or JPEG, one per request. So the first step is to rasterize the document, page by page. The simplest tool is pdftoppm, included with poppler-utils and available on all Linux distributions and via Homebrew on macOS.
Resolution deserves a minute of thought. DeepSeek-OCR works at fixed resolutions: 512, 640, 1024, or 1280 pixels per side depending on the mode, plus a dynamic mode called Gundam that divides the page into 640-pixel tiles around a global 1024 view. An A4 page rasterized at 200 dots per inch measures about 1 650 by 2 340 pixels; Ollama resizes it before sending it to the encoder. Increasing to 300 dots per inch adds nothing for the model and needlessly enlarges the files. Dropping to 100 loses the small characters in footnotes before the model even sees them.
If the scan is skewed or has very high contrast, the model generally handles it better than line-by-line OCR, but correcting the alignment beforehand is still beneficial. The site's Tesseract guide details useful preprocessing steps that apply here as well: https://quelllm.fr/guide/tesseract-ocr-guide-local
#Step 3: get Markdown with deepseek ocr
The official repository documents several prompts, each triggering different model behavior. The two main ones are “Convert the document to markdown.” for structured conversion, and “Free OCR.” for raw transcription without formatting. The prompts are in English in the training; keep them as is—the recognized text is output in the document's language. With the command-line client, place the image path directly in the prompt.
The same operation through the local API, which listens by default on http://localhost:11434, encodes the image as base64 in the images field. This is the preferred approach as soon as you process pages in sequence or call the model from another program.
Two options deserve to be pinned down. Setting the temperature to zero makes the output reproducible, which is what you expect from OCR. The 8,192-token context matches the model’s limit; below that, a dense page can be truncated in the middle of a table cell. The official repository also mentions dedicated instructions for describing a figure or locating text using coordinates; they are of little use for a simple PDF to index.
#Step 4: process an entire PDF
For a document several dozen pages long, a Python script a few lines long is enough. It processes the images in numerical order, calls Ollama page by page, removes location tags that the model may insert, and concatenates everything into a single Markdown file with one marker per page. The marker is useful later for finding the source page of a passage cited by your RAG.
Each call is independent: the model retains no memory of the previous page. That's a limitation for tables spanning two pages, and an advantage for robustness, since a failed page doesn't contaminate another. Ollama loads the model once and keeps it in memory between calls, according to the value of the OLLAMA_KEEP_ALIVE variable; you pay the loading time only at the start of the batch.
If you have a card with some headroom, the OLLAMA_NUM_PARALLEL option lets you process multiple pages at once, at the cost of additional VRAM for each context. On a 12 GB card, two parallel requests remain reasonable with a model this size; check with nvidia-smi that you are not spilling into system memory, because the slowdown is then dramatic.
#Step 5: clean up and validate the output
The Markdown produced is readable, but not necessarily suitable for indexing as-is. Before sending it to a vector database, review three points.
- 01Residual tagsWith the conversion prompt, the model is trained with a localization token and can return coordinates between ref and det tags. The script above removes them. Check that no image tag or prompt fragment remains at the beginning of the file either.
- 02Figures and tablesA vision-language model generates text, so it can invent a plausible value in an unreadable cell where a conventional OCR system would have left an aberrant character. Open financial or technical tables side by side with the scan and compare one line out of ten. For invoices, the site's dedicated guide explains how to cross-check against totals: https://quelllm.fr/guide/extraction-factures-ocr-llm
- 03Headers and footersPage numbers, document names, repeated legal notices: they appear on every page and pollute segment splitting. A regular expression targeting identical lines present on more than half the pages is usually enough to remove them.
- 04French accents and typographyThe publisher claims training in around a hundred languages. Still, check accented letters, French quotation marks, and nonbreaking spaces before double punctuation marks on a sample: that’s where subtle errors hide and later skew lexical search.
Once the file is clean, it enters a RAG pipeline like any other Markdown file: splitting by headings, embeddings, and a vector database. The site’s introduction to local RAG covers these steps: https://quelllm.fr/guide/rag-local-introduction
#Limitations and troubleshooting
- The model answers without looking at the image
- Either Ollama is too old for this architecture, or the image was not transmitted. On the command line, the path must be absolute or relative to the current directory, with no quotation marks around the path itself. Through the API, check that the images field contains valid base64 with no line breaks.
- Output cut off in the middle of a table
- The context is too short for the page. Set num_ctx to 8192 if you have not already. If the page is still too long, split it into two images, top and bottom, with an overlap of a few lines.
- Page read out of order
- With a three-column layout or a checkbox form, the model may mix up the blocks. Try the Free OCR prompt, which follows the spatial order more simply, or use Docling, whose layout analysis is explicit.
- Abnormally slow
- Use ollama ps to verify that the model is fully loaded on the GPU rather than partially on the CPU. An overly large context or a second loaded model can push it into system memory.
- Handwritten text
- DeepSeek-OCR is trained on printed documents and page renderings. Results vary widely with handwriting; do not rely on it without prior testing.
- Long documents and context across pages
- Each page is processed in isolation. A table that continues onto the next page loses its headers; you have to reinsert them manually or with a post-processing rule.
#When PaddleOCR, Docling, or Tesseract are enough
DeepSeek-OCR is not the answer for every scan. It excels when the page has a structure to preserve and you want Markdown without assembling a pipeline. In many common cases, a simpler or more specialized tool does just as well, using fewer resources and with less risk of fabrication.
- Tesseract
- Clean, printed text on a white background, in large quantities, when you only need the text. It runs on the CPU, generates nothing, and therefore invents nothing. Guide: https://quelllm.fr/guide/tesseract-ocr-guide-local
- PaddleOCR
- Text placed anywhere on the page, skewed scans, and tables that must be reconstructed with an explicit detection step before reading. Its VL variant is also a vision model, smaller than DeepSeek-OCR. Guide: https://quelllm.fr/guide/paddleocr-vl-ocr-local
- Docling
- Native PDFs, DOCX files, presentations: documents that already contain text and mainly require their layout and tables to be recovered. Docling can call an OCR engine for image-based pages, but its core is structural analysis. Guide: https://quelllm.fr/guide/docling-conversion-documents-ia
- DeepSeek-OCR
- Scans or photos of pages with headings, lists, tables, or formulas, plus the need for Markdown ready for indexing in a single pass on a modest graphics card.
One criterion often decides the matter: if an error in a number has consequences, choose a tool that recognizes without generating, or double-check with a second engine and compare the outputs. If the priority is making a heterogeneous document collection readable and searchable, DeepSeek-OCR's Markdown saves time at every subsequent step.
#Go further
- DeepSeek-OCR model sheet
- Parameters, VRAM by precision, license, and installation command. https://quelllm.fr/modele/deepseek-ocr
- Install Ollama
- Installation on Windows, macOS, and Linux, basic commands, and troubleshooting. https://quelllm.fr/guide/installer-ollama
- Local RAG without coding
- Connect the resulting Markdown to Open WebUI or AnythingLLM to query your documents. https://quelllm.fr/guide/rag-local-ollama-sans-coder
- Official sources
- GitHub repository deepseek-ai/DeepSeek-OCR (code, instructions, vLLM scripts for PDFs), Hugging Face page deepseek-ai/DeepSeek-OCR (weights and MIT license), Ollama page ollama.com/library/deepseek-ocr (tags, size, and minimum version).
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.