Intermediate 11 minSearch

Zotero + local LLM: summarize and query your bibliographie

Zotero centralizes your references and PDFs; a local LLM can read and synthesize them. By connecting the two, you get a research assistant capable of summarizing an article, comparing several papers, and preparing a literature review—without ever sending your unpublished work to a cloud. This guide shows how to build the complete Zotero + AI pipeline end to end, with Ollama and an RAG component.

By Clara M.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why connect a local AI to Zotero

A researcher quickly accumulates hundreds of papers in Zotero, most of them still unpublished, under embargo, or covered by a confidentiality clause. Pasting the PDF of a manuscript under review into a cloud chatbot means entrusting its contents to a third party—sometimes for training its models. Running an LLM locally solves the problem at its root: the files stay on your disk, inference runs on your machine, and nothing leaves.

The value of using Zotero as a research AI goes beyond privacy. Your library is already structured: collections by project, clean tags and metadata, annotated PDFs. It's a ready-to-use document repository for an LLM. Instead of downloading and reorganizing papers again, you connect the model directly to what Zotero has already organized.

Summarize
Get the objective, method, and results of a twenty-page article in your language in thirty seconds.
Query
Ask a question and let the model search for the answer across an entire collection, not just a single PDF.
Compare
Compare the methods or results from several papers to sketch out the state of the art.
Privacy
Process manuscripts under embargo or collaborator data without a data-leakage clause.
i
What local setups cannot replace
A local LLM saves you reading time, not rigor. It produces draft summaries to review, not truths to quote as-is. The section on limitations explains where it still gets things wrong.

#How Zotero and the LLM communicate

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Zotero does not natively “talk” to a language model. The bridge is the files: Zotero stores each PDF on your disk, and that PDF is what you give the LLM to read. Two main approaches coexist, and they are often combined.

On-demand summary
We extract the text from a specific PDF and send it to Ollama with a summarization instruction. Simple, fast, and ideal for one article at a time.
RAG over the library
You index an entire folder of PDFs in a vector database, then query it in natural language. The model reads only the relevant passages, allowing you to cover dozens of papers.

In both cases, the central piece remains Zotero's storage folder, where all the PDFs reside. Zotero 7 also exposes a local API (on port 23119) that lets you list its items programmatically, but to get started, pointing directly to the files is the most robust and transparent approach.

#Prerequisites

Zotero 7
With your PDFs attached to the references (not just the metadata). The built-in PDF reader is not required, but the files must be stored locally.
Ollama
The daemon that serves models, on http://localhost:11434 by default. See the installation guide if you haven't done so yet.
A summarization model
Qwen 3.5 9B (≈6.6 GB of VRAM in Q4_K_M, 256k context, and vision) or Mistral Small 24B for excellent French; a Qwen 3.5 4B (3.4 GB) is sufficient on a small setup.
An embeddings model
For RAG: nomic-embed-text or mxbai-embed-large, retrievable through ollama pull. They are lightweight and run even without a GPU.
An optional RAG component
AnythingLLM or Open WebUI if you want to query your entire bibliography without coding.
Retrieve the models
# Modèle de synthèse (adaptez à votre VRAM)
ollama pull qwen3.5:9b

# Modèle d'embeddings pour le RAG
ollama pull nomic-embed-text
→
Choose the size based on the GPU
With Q4_K_M, plan on approximately 5 GB of VRAM for a 7B, 9 GB for a 14B, and 19 GB for a 32B. A RTX 3060 12 GB can run a 14B comfortably; on a Mac with unified memory, a 32B remains feasible.

#Step 1: Find your PDFs in Zotero

Zotero stores each attachment in a subdirectory of the storage directory, named with an eight-character key. You don't need to understand this directory structure in detail: you just need to know where it is so you can point the LLM to it.

Windows
C:\Users\<vous>\Zotero\storage
macOS
~/Zotero/storage
Linux
~/Zotero/storage

If in doubt, the exact path is shown in Zotero under Edit ▸ Settings ▸ Advanced ▸ Files and Folders ▸ “Data Directory.” The storage folder is directly inside it.

To work cleanly on a specific project, it is best to export a collection rather than search through all of storage. Right-click a collection ▸ “Export Collection”: Zotero can copy the PDFs and a bibliography (BibTeX, CSV, JSON) into a dedicated folder, which you can then provide to the LLM.

i
Extract the text, not the image
An LLM reads text, not pixels. A “scanned” PDF without a text layer must first go through OCR (ocrmypdf, for example). PDFs from modern publishers already contain a text layer that can be used directly.

#Step 2: summarize a scientific article

The most common case: you open a paper in Zotero and want its summary before deciding whether it deserves a full read. We extract the text from the PDF, then send it to Ollama. The pymupdf (fitz) library extracts it cleanly and quickly.

Dependencies
pip install pymupdf requests
resumer_article.py
import sys, fitz, requests

def extraire_texte(pdf_path):
    doc = fitz.open(pdf_path)
    return "\n".join(page.get_text() for page in doc)

SYSTEM = (
    "Tu es un assistant de recherche. Tu résumes des articles "
    "scientifiques en français, avec rigueur, sans rien inventer. "
    "Tu t'appuies uniquement sur le texte fourni."
)

texte = extraire_texte(sys.argv[1])

resp = requests.post("http://localhost:11434/api/chat", json={
    "model": "qwen3.5:9b",
    "stream": False,
    "options": {"num_ctx": 16384, "temperature": 0.2},
    "messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": PROMPT + "\n\n---\n" + texte},
    ],
})

print(resp.json()["message"]["content"])

We call the Ollama REST API on port 11434: it separates the system message (the role) from the content (the article). Two settings matter here — a low temperature (0.2) to limit embellishment, and a sufficiently large num_ctx to ingest an entire paper without truncating it.

!
The context window is your hard limit
A twenty-page article often contains 12 000 to 18 000 words. If num_ctx is too small, the model will see only the beginning of the PDF and blindly summarize the end. Check the model's context, increase num_ctx if VRAM allows, or split the document into sections.

#A reliable summarization prompt

Summary quality depends more on the prompt than on the model. A vague prompt produces a generic paragraph; a structured prompt produces a reusable reading brief. The secret is to require a section-based format modeled on the structure of a scientific paper and forbid the model from going beyond the text.

Reading-sheet prompt
Rédige une fiche de lecture en français de l'article ci-dessous, en respectant EXACTEMENT ce format :

## Question de recherche
Ce que l'article cherche à établir, en une à deux phrases.

## Méthode
Données, protocole, modèles ou outils utilisés.

## Résultats principaux
- Une puce par résultat marquant, chiffré si l'article donne des chiffres.

## Limites annoncées
- Les limites que les auteurs reconnaissent eux-mêmes.

## À retenir pour mon état de l'art
Deux à trois phrases sur l'apport et le positionnement de l'article.

Règles :
- Ne reprends que ce qui est réellement écrit dans l'article.
- Si une section n'est pas renseignée dans le texte, écris « non précisé ».
- N'invente aucun chiffre, aucune référence, aucun nom d'auteur.
The imposed format
Section headings force the model to sort instead of rambling. You get the same framework from one article to the next, comparable at a glance.
The anti-hallucination rule
“Only repeat what is written” and “not specified” reduce the risk that the LLM will invent a result or reference—the main danger in scientific contexts.
The state-of-the-art section
By explicitly requesting the positioning, you turn the summary into raw material that can be directly reused in a literature review.
→
Always verify the figures at the source
Even with a low temperature and a strict prompt, an LLM can transpose a number from one line to another. Treat every number in the summary as a lead to confirm in the PDF before citing it.

#Step 3: query your entire bibliography with RAG

Summarizing a PDF is useful; querying fifty papers at once is what changes how you prepare a literature review. That is where RAG (Retrieval-Augmented Generation) comes in: you split PDFs into chunks, convert them into vectors with the embedding model, and the LLM reads only the passages relevant to each question.

The simplest approach, without writing code, is to point AnythingLLM or Open WebUI to the folder exported from Zotero (or directly to storage). These tools handle indexing and the vector database for you; you only need to choose Ollama as the provider and nomic-embed-text as the embeddings model.

  1. 01
    Export the collection
    From Zotero, export the relevant collection with its PDFs to a dedicated folder, such as ~/recherche/etat-de-l'art.
  2. 02
    Create a workspace
    In AnythingLLM, create a workspace and set the LLM provider to Ollama (http://localhost:11434) and the embedding model to nomic-embed-text.
  3. 03
    Import documents
    Drag the PDF folder into the workspace. The tool automatically extracts, chunks, and indexes the text.
  4. 04
    Query in natural language
    Ask cross-cutting questions: “Which evaluation methods come up most often?”, “Which papers contradict hypothesis X?”

RAG’s main advantage over simple summarization: the model cites the passages it used. You can trace them back to the source PDF to verify them, which is essential in research. Always ask the model to indicate which document each claim comes from.

i
RAG does not read “everything”
Contrary to a common assumption, RAG does not make the model read all the papers: it retrieves only the few passages closest to the question. A poorly worded question retrieves the wrong excerpts. Rephrase it if the answer seems off-topic.

#The limitations on highly technical papers

This is where you need to be honest: in a math, theoretical physics, or highly formal ML article, a local LLM quickly shows its limitations. Understanding these blind spots helps you avoid trusting it in the wrong place.

The formulas
Text extraction flattens equations: subscripts, exponents, and Greek symbols get mixed up or disappear. The model then reasons over corrupted formulas and may draw false conclusions.
The tables
A table turns into a jumble of numbers when extracted as plain text. Row/column relationships are lost, so figures cited from a table are unreliable.
The figures
A text model cannot see charts. Anything that exists only in a figure is completely lost to it—it won't mention it, or it will make it up from the caption.
Mathematical reasoning
Following a demonstration or checking a derivation exceeds what a local 14B model can do reliably. It paraphrases the proof without validating it.

In practice: use the LLM for the “narrative” layer of a paper — research question, motivation, stated contributions, limitations, and positioning. Keep human oversight for every formula, numerical table, and figure. For papers rich in equations, a multimodal model that can read pages as images rather than extracted text delivers better results, but it remains a supplement, not a substitute for your own reading.

!
Never cite blindly
Never copy a number, formula, or claim produced by the model without finding it again in the PDF. The LLM’s role is to guide you and do the initial legwork, not replace scientific verification.

#Troubleshooting

The PDF comes out blank
This is a scan without a text layer. Run it through OCR (ocrmypdf entree.pdf sortie.pdf) before extraction.
The summary skips the end of the article
num_ctx is too small: the text was truncated. Increase it if VRAM allows; otherwise, summarize by section and then summarize the summaries.
RAG gives irrelevant answers
The wrong passages were retrieved. Rephrase the question using the domain's exact terms, or reduce chunk size during indexing.
Slow responses
A 9B model crawls on the CPU. Make sure Ollama is actually using the GPU, or switch to a Qwen 3.5 4B (3.4 GB) for routine summarization.
Numbers that don't match the PDF
A classic symptom of a poorly extracted table or a hallucination. Lower the temperature and, above all, cross-check against the source.

#Go further

This guide builds on components already covered in detail on the site. To explore installation or RAG in particular:

Install Ollama: Windows, macOS, and Linux
The starting point for serving your models locally on port 11434, if you haven't set it up yet.
Local RAG with Ollama without coding (Open WebUI, AnythingLLM)
The definitive guide to building the RAG component that queries your entire bibliography without writing a line of code.
Local RAG with LM Studio: chat with your documents
An all-in-one alternative if you prefer LM Studio to the Ollama + separate interface pair.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.