Intermediate 11 minDocuments

DeepSeek-OCR locally: read its PDFs scanned

DeepSeek-OCR is a vision model with about 3 billion parameters, released under an MIT license and designed to turn an image of a page into structured text, including Markdown. This guide shows how to run deepseek ocr locally with Ollama: which runtime version to use, how much memory to plan for, how to split a scanned PDF into pages, and how to retrieve Markdown usable by a RAG. It ends with cases where a simpler tool does the job better.

By Léa B.·Update 2026-10-07·Tested on Windows, macOS, and Linux

#Why DeepSeek-OCR instead of traditional OCR

A scanned PDF contains no text, only images of text. A traditional OCR engine such as Tesseract extracts raw lines from it: it doesn't know that a block is a heading, it breaks tables, and it loses the reading order of a two-column page. DeepSeek-OCR approaches the problem differently. It's a vision-language model: it looks at the entire page, then directly generates Markdown with its headings, lists, tables, and formulas.

The model was released by DeepSeek in fall 2025 with a paper titled “Contexts Optical Compression.” The central idea is to represent a page with a small number of visual tokens, between 64 and 400 depending on the resolution mode, instead of the thousands of text tokens it would contain. The publisher claims decoding accuracy of approximately 97% when the compression ratio remains below ten. That is its figure, measured on its own datasets, not ours: the main takeaway is that the model is lightweight and fast for what it does.

For local use, three things matter. The model fits on an entry-level graphics card. It outputs Markdown that can be indexed as-is in a RAG pipeline. And it is available in the Ollama library, avoiding the need to install PyTorch, Flash Attention, and vLLM as required by the official repository. Full specifications are on the model page: https://quelllm.fr/modele/deepseek-ocr

#Prerequisites and memory requirements

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The site’s model page provides VRAM guidelines by precision: approximately 2 GB in Q4, 4 GB in Q8, and 6 GB in FP16, for a context of 8,192 tokens. The actual download size depends on what Ollama packaged into the tag; the library’s tag page displays the file size, and that is the authoritative figure. Also account for context memory: a dense page converted to Markdown can produce several thousand output tokens.

GPU NVIDIA
Any card with 8 GB or more is comfortable. A RTX 3060 12 GB or a RTX 4070 12 GB can run the model in FP16 with room for context.
Mac Apple Silicon
Ollama uses unified memory. A Mac with 16 GB is enough for the model alone; plan on 24 GB if you want to keep a chat model loaded alongside it for RAG.
Without a GPU
Possible, but image encoding and Markdown generation run on the processor. For a few pages, it’s tolerable; for a batch of one hundred pages, expect minutes per page depending on the machine.
Software
Ollama up to date, poppler-utils for splitting PDFs, and Python 3 with the ollama library if you want to automate batch processing.
i
This guide measures nothing
No throughput in pages per minute or accuracy score is announced here. The figures cited are those from DeepSeek's publication. Quality on your documents depends on scan resolution, language, and layout: test on twenty representative pages before processing a batch.

#Step 1: verify Ollama and retrieve the correct tag

DeepSeek-OCR arrived in the Ollama library in late 2025, after the model was released. It relies on a specialized vision encoder, DeepEncoder, which combines a local-window module with a global-attention module. Ollama had to add support for this architecture to its engine: an earlier runtime downloads the weights but refuses to load them, with a message saying that the model requires a newer version. The library page displays the minimum required version when applicable.

Terminal
# Version installée (comparez avec la version minimale affichée sur ollama.com/library/deepseek-ocr)
ollama -v

# Récupérer le modèle (le tag 3b est celui référencé par la fiche modèle)
ollama pull deepseek-ocr:3b

# Vérifier l'architecture, le contexte et la quantification empaquetée
ollama show deepseek-ocr:3b

The show command is the only reliable source for knowing what you just downloaded: it reports the model family, parameter count, context length, and quantization level. If the latest tag and the 3b tag point to the same digest, it makes no difference which one you use; the guide uses 3b to remain explicit.

!
Update Ollama before looking any further
If the model downloads but fails to load, or if Ollama responds with text without taking the image into account, the most common cause is an outdated runtime. On Linux, rerun the official installation script; on Windows and macOS, the application updates itself or through the menu. Then run ollama -v again.

#Step 2: prepare the PDF pages

Ollama doesn't open PDFs. It accepts images, PNG or JPEG, one per request. So the first step is to rasterize the document, page by page. The simplest tool is pdftoppm, included with poppler-utils and available on all Linux distributions and via Homebrew on macOS.

Terminal
# Debian/Ubuntu : sudo apt install poppler-utils
# macOS : brew install poppler
# Arch : sudo pacman -S poppler

mkdir -p pages
pdftoppm -r 200 -png scan.pdf pages/page
ls pages
# pages/page-1.png  pages/page-2.png  ...  (numérotation complétée par des zéros selon le nombre de pages)

Resolution deserves a minute of thought. DeepSeek-OCR works at fixed resolutions: 512, 640, 1024, or 1280 pixels per side depending on the mode, plus a dynamic mode called Gundam that divides the page into 640-pixel tiles around a global 1024 view. An A4 page rasterized at 200 dots per inch measures about 1 650 by 2 340 pixels; Ollama resizes it before sending it to the encoder. Increasing to 300 dots per inch adds nothing for the model and needlessly enlarges the files. Dropping to 100 loses the small characters in footnotes before the model even sees them.

If the scan is skewed or has very high contrast, the model generally handles it better than line-by-line OCR, but correcting the alignment beforehand is still beneficial. The site's Tesseract guide details useful preprocessing steps that apply here as well: https://quelllm.fr/guide/tesseract-ocr-guide-local

#Step 3: get Markdown with deepseek ocr

The official repository documents several prompts, each triggering different model behavior. The two main ones are “Convert the document to markdown.” for structured conversion, and “Free OCR.” for raw transcription without formatting. The prompts are in English in the training; keep them as is—the recognized text is output in the document's language. With the command-line client, place the image path directly in the prompt.

Terminal
# Conversion structurée (titres, listes, tableaux)
ollama run deepseek-ocr:3b "Convert the document to markdown. ./pages/page-1.png"

# Transcription brute, sans structure
ollama run deepseek-ocr:3b "Free OCR. ./pages/page-1.png"

The same operation through the local API, which listens by default on http://localhost:11434, encodes the image as base64 in the images field. This is the preferred approach as soon as you process pages in sequence or call the model from another program.

Terminal
# Linux (GNU coreutils). Sur macOS, remplacez base64 -w0 par base64 -i pages/page-1.png
curl http://localhost:11434/api/generate -d "{
  \"model\": \"deepseek-ocr:3b\",
  \"prompt\": \"Convert the document to markdown.\",
  \"images\": [\"$(base64 -w0 pages/page-1.png)\"],
  \"stream\": false,
  \"options\": { \"num_ctx\": 8192, \"temperature\": 0 }
}"

Two options deserve to be pinned down. Setting the temperature to zero makes the output reproducible, which is what you expect from OCR. The 8,192-token context matches the model’s limit; below that, a dense page can be truncated in the middle of a table cell. The official repository also mentions dedicated instructions for describing a figure or locating text using coordinates; they are of little use for a simple PDF to index.

→
Test both prompts on the same page
On a page of regular text, Free OCR often produces a cleaner and faster result. On a page with a table of figures or multiple heading levels, Markdown conversion justifies its overhead. Choose by document type, not once and for all.

#Step 4: process an entire PDF

For a document several dozen pages long, a Python script a few lines long is enough. It processes the images in numerical order, calls Ollama page by page, removes location tags that the model may insert, and concatenates everything into a single Markdown file with one marker per page. The marker is useful later for finding the source page of a passage cited by your RAG.

Terminal
pip install ollama
ocr_pdf.py
import pathlib
import re
import ollama

MODEL = "deepseek-ocr:3b"
PROMPT = "Convert the document to markdown."

# Tri numérique : page-2 avant page-10
pages = sorted(
    pathlib.Path("pages").glob("page-*.png"),
    key=lambda p: int(re.search(r"(\d+)", p.stem).group(1)),
)

parts = []
for page in pages:
    r = ollama.generate(
        model=MODEL,
        prompt=PROMPT,
        images=[str(page)],
        options={"num_ctx": 8192, "temperature": 0},
    )
    md = r["response"]
    # Balises de localisation éventuelles : on garde le texte, on jette les coordonnées
    md = re.sub(r"<\|ref\|>(.*?)<\|/ref\|><\|det\|>.*?<\|/det\|>", r"\1", md, flags=re.S)
    parts.append(f"<!-- {page.name} -->\n{md.strip()}\n")
    print(f"{page.name} : {len(md)} caractères")

pathlib.Path("scan.md").write_text("\n\n".join(parts), encoding="utf-8")

Each call is independent: the model retains no memory of the previous page. That's a limitation for tables spanning two pages, and an advantage for robustness, since a failed page doesn't contaminate another. Ollama loads the model once and keeps it in memory between calls, according to the value of the OLLAMA_KEEP_ALIVE variable; you pay the loading time only at the start of the batch.

If you have a card with some headroom, the OLLAMA_NUM_PARALLEL option lets you process multiple pages at once, at the cost of additional VRAM for each context. On a 12 GB card, two parallel requests remain reasonable with a model this size; check with nvidia-smi that you are not spilling into system memory, because the slowdown is then dramatic.

#Step 5: clean up and validate the output

The Markdown produced is readable, but not necessarily suitable for indexing as-is. Before sending it to a vector database, review three points.

  1. 01
    Residual tags
    With the conversion prompt, the model is trained with a localization token and can return coordinates between ref and det tags. The script above removes them. Check that no image tag or prompt fragment remains at the beginning of the file either.
  2. 02
    Figures and tables
    A vision-language model generates text, so it can invent a plausible value in an unreadable cell where a conventional OCR system would have left an aberrant character. Open financial or technical tables side by side with the scan and compare one line out of ten. For invoices, the site's dedicated guide explains how to cross-check against totals: https://quelllm.fr/guide/extraction-factures-ocr-llm
  3. 03
    Headers and footers
    Page numbers, document names, repeated legal notices: they appear on every page and pollute segment splitting. A regular expression targeting identical lines present on more than half the pages is usually enough to remove them.
  4. 04
    French accents and typography
    The publisher claims training in around a hundred languages. Still, check accented letters, French quotation marks, and nonbreaking spaces before double punctuation marks on a sample: that’s where subtle errors hide and later skew lexical search.

Once the file is clean, it enters a RAG pipeline like any other Markdown file: splitting by headings, embeddings, and a vector database. The site’s introduction to local RAG covers these steps: https://quelllm.fr/guide/rag-local-introduction

#Limitations and troubleshooting

The model answers without looking at the image
Either Ollama is too old for this architecture, or the image was not transmitted. On the command line, the path must be absolute or relative to the current directory, with no quotation marks around the path itself. Through the API, check that the images field contains valid base64 with no line breaks.
Output cut off in the middle of a table
The context is too short for the page. Set num_ctx to 8192 if you have not already. If the page is still too long, split it into two images, top and bottom, with an overlap of a few lines.
Page read out of order
With a three-column layout or a checkbox form, the model may mix up the blocks. Try the Free OCR prompt, which follows the spatial order more simply, or use Docling, whose layout analysis is explicit.
Abnormally slow
Use ollama ps to verify that the model is fully loaded on the GPU rather than partially on the CPU. An overly large context or a second loaded model can push it into system memory.
Handwritten text
DeepSeek-OCR is trained on printed documents and page renderings. Results vary widely with handwriting; do not rely on it without prior testing.
Long documents and context across pages
Each page is processed in isolation. A table that continues onto the next page loses its headers; you have to reinsert them manually or with a post-processing rule.
i
A second version exists
DeepSeek published a second version of the model in early 2026, DeepSeek-OCR-2, with a redesigned encoder. As of writing, we were unable to confirm its presence in the Ollama library. Check the publisher's Hugging Face page and the site's model card before choosing; the method described here remains the same.

#When PaddleOCR, Docling, or Tesseract are enough

DeepSeek-OCR is not the answer for every scan. It excels when the page has a structure to preserve and you want Markdown without assembling a pipeline. In many common cases, a simpler or more specialized tool does just as well, using fewer resources and with less risk of fabrication.

Tesseract
Clean, printed text on a white background, in large quantities, when you only need the text. It runs on the CPU, generates nothing, and therefore invents nothing. Guide: https://quelllm.fr/guide/tesseract-ocr-guide-local
PaddleOCR
Text placed anywhere on the page, skewed scans, and tables that must be reconstructed with an explicit detection step before reading. Its VL variant is also a vision model, smaller than DeepSeek-OCR. Guide: https://quelllm.fr/guide/paddleocr-vl-ocr-local
Docling
Native PDFs, DOCX files, presentations: documents that already contain text and mainly require their layout and tables to be recovered. Docling can call an OCR engine for image-based pages, but its core is structural analysis. Guide: https://quelllm.fr/guide/docling-conversion-documents-ia
DeepSeek-OCR
Scans or photos of pages with headings, lists, tables, or formulas, plus the need for Markdown ready for indexing in a single pass on a modest graphics card.

One criterion often decides the matter: if an error in a number has consequences, choose a tool that recognizes without generating, or double-check with a second engine and compare the outputs. If the priority is making a heterogeneous document collection readable and searchable, DeepSeek-OCR's Markdown saves time at every subsequent step.

#Go further

DeepSeek-OCR model sheet
Parameters, VRAM by precision, license, and installation command. https://quelllm.fr/modele/deepseek-ocr
Install Ollama
Installation on Windows, macOS, and Linux, basic commands, and troubleshooting. https://quelllm.fr/guide/installer-ollama
Local RAG without coding
Connect the resulting Markdown to Open WebUI or AnythingLLM to query your documents. https://quelllm.fr/guide/rag-local-ollama-sans-coder
Official sources
GitHub repository deepseek-ai/DeepSeek-OCR (code, instructions, vLLM scripts for PDFs), Hugging Face page deepseek-ai/DeepSeek-OCR (weights and MIT license), Ollama page ollama.com/library/deepseek-ocr (tags, size, and minimum version).
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.