Intermediate 11 minConcepts

Karpathy’s LLM Wiki: a knowledge base locale

LLM Wiki is a pattern described by Andrej Karpathy in April 2026: instead of searching through your documents for every question, a model writes and maintains a Markdown wiki that grows with each new source. This guide explains the principle based on his post and gist, proposes a local setup with Ollama and a terminal agent, then clarifies what this pattern does not replace in RAG. It contains neither an in-house test nor a numerical comparison: statements about the pattern are attributed to their source, while the rest is identified as our implementation.

By Clara M.·Update 2026-10-06·Tested on Windows, macOS, and Linux

#LLM Wiki: the principle in two minutes

On April 2, 2026, Andrej Karpathy described on X how he uses language models to build personal knowledge bases. Two days later, he published a gist titled “LLM Wiki,” which he presented as a pattern for building this type of database with an LLM. Both texts are short and take ten minutes to read; the links appear at the end of the page.

The gist starts with an observation: most use cases combining LLMs and documents look like RAG. You upload files, the system retrieves excerpts when you ask a question, and the model writes a response. Karpathy acknowledges that this works, but points out that the model rediscovers the knowledge with every question and nothing accumulates. A question requiring five documents to be cross-referenced means finding and piecing together the same parts every time.

The LLM Wiki moves the work upstream. When a new source arrives, the model doesn't just index it: it reads it, extracts the essentials, and integrates them into a set of interconnected Markdown pages. It updates existing pages, revises summaries, and notes where the new source contradicts what was written. The gist describes a persistent artifact that improves over time: the cross-references are already done when the question arrives.

The division of roles is clear in the text. The human chooses the sources, explores, and asks the questions. The model does everything else: summarizing, connecting, classifying, and maintaining the records. Karpathy says he works with the open agent on one side of the screen and Obsidian on the other, and sums up the setup with an image: Obsidian is the IDE, the LLM is the programmer, and the wiki is the codebase.

i
An idea, not software
The gist is presented as an “idea file” to copy and paste into your own agent (it mentions OpenAI Codex, Claude Code, OpenCode, or Pi), which will work out the details with you. There is therefore no official repository to clone or version to install. The practical part of this guide is one possible implementation among others, not a reference.

#Three layers: sources, wiki, conventions

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The gist describes a three-layer architecture. None requires any particular tool: they are directories and text files.

The raw sources
Your document collection: articles, research papers, images, and data files. They are immutable: the model reads them and never modifies them. They—not the wiki—are authoritative.
The wiki
A folder of Markdown files written by the model: source summaries, entity pages, concept pages, comparisons, and an overall synthesis. This layer belongs to the model, which creates the pages, updates them, and maintains the links. You read; it writes.
The conventions file
A document that tells the model how the wiki is structured, which rules to follow, and how to ingest a source, answer a question, or clean things up. The gist cites CLAUDE.md for Claude Code and AGENTS.md for Codex. This file is what turns a general-purpose agent into a disciplined wiki maintainer.

Two special files help the model—and you—find your way around. The first, index.md, is a catalog: every page appears there with a link and a one-line summary, organized by category. To answer a question, the model first reads the index, then opens the relevant pages. The second, log.md, is a chronological, append-only log: ingestions, questions, and verification passes.

The gist suggests starting each log entry with a consistent prefix, which makes it possible to filter the log with simple Unix tools. The example has this form:

Suggested log input format in the gist
## [2026-04-02] ingest | Article Title

#Three operations: ingest, query, verify

Ingest
You drop a source into the raw sources folder and ask the model to process it. According to the gist, it reads the source, discusses the key points with you, writes a summary page, updates the index as well as the relevant entity and concept pages, then adds an entry to the log. Karpathy says a single source can affect 10 to 15 wiki pages.
Query
You ask a question. The model searches for the relevant pages, reads them, and writes an answer that cites its references. The gist emphasizes one point: a good answer can be filed in the wiki as a new page, so your explorations accumulate instead of disappearing into a discussion’s history.
Check (lint)
From time to time, you ask the model to perform a wiki health check: contradictions between pages, claims superseded by newer sources, orphaned pages with no inbound links, concepts mentioned without a dedicated page, and missing cross-references.

Why entrust this work to a model? The gist's argument is simple: what kills personal wikis is neither reading nor thinking, but record-keeping. Updating cross-references, keeping summaries current, flagging contradictions: the maintenance burden grows faster than the wiki's value, and people give up. A model doesn't get tired and can modify fifteen files in one pass. Karpathy links the idea to the Memex envisioned by Vannevar Bush in 1945.


#Requirements for a wiki maintained by a local model

The gist assumes no specific provider. It needs an agent capable of reading and writing files, driven by a model. Locally, this yields the following building blocks.

Ollama, up to date
It serves the model on http://localhost:11434. The ollama launch command used below exists only in recent versions. Installation is covered in our “Install Ollama” guide.
An agent with file access
A chat interface is not enough: you need a tool that opens, creates, and modifies files on disk. The example below uses OpenCode, an open-source terminal agent mentioned in the gist, which reads an AGENTS.md file placed at the root of the directory.
A model that can call tools
Reading and writing files happens through tool calls. Choose a model that lists tools capability in the Ollama library. Our “OpenCode + Ollama” guide lists several, including qwen3-coder:30b, devstral-small-2:24b, and gpt-oss:20b.
Memory for context
Q4_K_M reference figures for weights alone: about 5 GB for 7 billion parameters, 9 GB for 14 billion, and 19 GB for 32 billion. The Ollama documentation requires at least 64,000 tokens of context for agents, which must be added to these figures.
Git
The wiki is just a folder of Markdown files: the gist points out that turning it into a Git repository provides version history without adding anything.
Obsidian (optional)
To read the wiki, follow links, and display the page graph. Any Markdown editor will do; Obsidian is not involved in writing.

#Step-by-step setup with Ollama

The commands are written for macOS and Linux; on Windows, the simplest option is to use WSL. The directory tree and conventions file are examples to adapt: the gist specifies that the folder structure, conventions, and page format depend on your domain and model, and that everything there is optional and modular.

  1. 01
    Create the folder and repository
    One folder for raw sources, one folder for pages, an index, and a log, all under git.
  2. 02
    Write the conventions file
    An AGENTS.md at the root describing the structure, writing rules, and the three procedures: ingestion, querying, and verification.
  3. 03
    Launch the model and agent
    Ollama serves a tool-capable model with 64,000 tokens of context; OpenCode opens in the wiki folder.
  4. 04
    Ingest a first source
    A single document, processed before your eyes, reviewed, then committed to git.
  5. 05
    Query and organize the answers
    Ask questions in the wiki; useful summaries become pages.
  6. 06
    Check regularly
    A review pass that lists contradictions, orphaned pages, and broken links.

#1. Create the folder and repository

Terminal
mkdir -p ~/wiki/raw
mkdir -p ~/wiki/wiki/sources ~/wiki/wiki/entites ~/wiki/wiki/concepts
cd ~/wiki
touch wiki/index.md wiki/log.md
git init

The raw/ folder will receive your documents, and the wiki/ folder the pages written by the model. Names are up to you, as long as the separation between the two remains clear.

#2. Write the conventions file

This is the part that matters most. Without it, the agent improvises a different structure in every session. Create an AGENTS.md file at the root of ~/wiki, for example based on this:

~/wiki/AGENTS.md
# Conventions du wiki

Tu es le mainteneur de ce wiki. Tu écris et tu mets à jour les pages.
L'humain choisit les sources et pose les questions.

## Structure
- raw/ : sources brutes. Lecture seule : ne jamais modifier, renommer ni supprimer.
- wiki/sources/ : une page de résumé par source.
- wiki/entites/ : une page par personne, organisation, outil ou produit.
- wiki/concepts/ : une page par notion.
- wiki/index.md : catalogue de toutes les pages (lien + résumé d'une ligne), par catégorie.
- wiki/log.md : journal chronologique, ajout seul.

## Règles d'écriture
- Noms de fichiers en minuscules, avec tirets, sans accents.
- Liens internes au format [[nom-de-page]].
- Chaque affirmation renvoie au fichier de raw/ dont elle vient.
- Si deux sources se contredisent, garder les deux versions et le signaler.
- Avant de créer une page, vérifier dans l'index qu'elle n'existe pas déjà.

## Ingestion
Quand on te demande d'ingérer un fichier de raw/ :
1. Lire la source en entier.
2. Présenter les points clés et attendre la validation.
3. Écrire la page de résumé dans wiki/sources/.
4. Mettre à jour ou créer les pages d'entités et de concepts concernées.
5. Mettre à jour wiki/index.md.
6. Ajouter une entrée à wiki/log.md : ## [AAAA-MM-JJ] ingest | Titre

## Question
1. Lire wiki/index.md, puis les pages utiles.
2. Répondre en citant les pages et les sources brutes.
3. Ne rien affirmer qui ne figure pas dans le wiki ; dire ce qui manque.
4. Si la réponse apporte une synthèse nouvelle, proposer de l'enregistrer comme page.

## Vérification
Signaler sans corriger d'office : contradictions, affirmations dépassées,
pages orphelines, concepts cités sans page, liens cassés.

This file is ours, not Karpathy's: the gist describes the role of the conventions file without providing a template, and recommends evolving it with the model as you learn what works in your domain. Keep it short. The agent rereads it every session, and every line takes up context space.

#3. Launch the model and agent

Download a model capable of calling tools, then open OpenCode in the wiki directory. The command ollama launch opencode starts OpenCode with a model served by Ollama, selected from the picker. OpenCode installation itself is described in our “OpenCode + Ollama” guide.

Terminal
# Un modèle généraliste avec appel d'outils (à adapter à votre mémoire)
ollama pull gpt-oss:20b

# Ouvrir l'agent dans le dossier du wiki
cd ~/wiki
ollama launch opencode

The remaining issue is context. According to the Ollama documentation, the default window depends on VRAM (4,000 tokens below 24 GB, 32,000 between 24 and 48 GB), while agents require at least 64,000. The OLLAMA_CONTEXT_LENGTH variable sets it when the server starts; if Ollama is already running as an application or service, set the value in its settings instead of starting a second server.

Terminal — Ollama server with 64,000 context tokens
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

# Dans un autre terminal, une fois le modèle chargé :
ollama ps

The ollama ps command indicates whether the model fits entirely on the GPU. If it spills over to the processor, every ingestion becomes very slow: choose a smaller model before cutting back on context.

#4. Ingest a first source

Drop an initial document into raw/, preferably in Markdown or plain text. For web pages, the gist points to the Obsidian Web Clipper extension, which converts an article into a Markdown file. Then give the agent the instruction:

Instruction to the agent
Ingère raw/mon-premier-article.md en suivant AGENTS.md.
Présente-moi d'abord les points clés et attends ma validation avant d'écrire.

The agent reads the source, proposes its key points, then creates and edits the pages. Review the result before going any further: the summary page, the created entity pages, the index, and the log. Then save the wiki state.

Terminal
git status
git add -A
git commit -m "ingest: mon-premier-article"
→
One source at a time, one commit at a time
Karpathy says he prefers ingesting sources one at a time while staying involved, rather than processing them in batches. With a local model, it’s also a matter of context: one source at a time leaves room for pages to review and modify. Committing after each ingestion gives you a rollback point: git diff shows exactly what the agent changed, and returning to the previous commit undoes a failed ingestion.

#5. Query and organize the responses

After adding a few sources, ask the agent your questions in the same folder. Explicitly ask it to cite its pages and sources, and to say what the wiki does not contain.

Instruction to the agent
D'après le wiki, qu'est-ce qui distingue l'approche A de l'approche B ?
Cite les pages et les sources brutes utilisées, et signale ce qui manque.
Si la réponse apporte une synthèse nouvelle, enregistre-la dans wiki/concepts/
puis mets à jour l'index et le journal.

#6. Check regularly

Every few ingestion runs, perform a verification pass. Ask for a list of issues before making any corrections: you stay in control of what is merged, renamed, or deleted.

Instruction to the agent
Fais une passe de vérification du wiki en suivant AGENTS.md.
Liste les problèmes trouvés, sans rien modifier pour l'instant.

You can inspect the log without an agent. With the regular entry prefix, the command provided in the gist displays the latest operations (only the path is adapted to our directory structure):

Terminal — last five log entries
grep "^## \[" wiki/log.md | tail -5

#LLM Wiki or RAG: what the prompt cannot replace

The gist contrasts a wiki with RAG to explain the idea. It does not say that one replaces the other, and this guide does not say that either: the two approaches address different situations. Here is what separates them, without numbers, because we have no measurements to present.

The work moment
A RAG works question by question: it searches for excerpts, then the model writes the response. A wiki works at ingestion time: the summary is written once, then reread for each question.
What is retained
A RAG stores excerpts and their vectors, unreadable as-is. A wiki stores written pages that you can read, edit, and version.
L'infrastructure
A RAG system requires an embeddings model, a vector database, and a chunking strategy. The wiki requires a folder and an agent. According to the gist, the index is sufficient at a moderate scale (on the order of a hundred sources and a few hundred pages) and avoids setting up an embeddings-based RAG infrastructure.
Source fidelity
A RAG gives the model original passages. The wiki gives it a reformulation written by a model, with the risk of error that entails.
Volume
A RAG is designed for large corpora. The wiki is bounded by what the model can read at once: the index, relevant pages, and source must fit within the context window.

Beyond a moderate scale, the gist brings search back in. It cites qmd, a local search engine for Markdown files that combines BM25, vector search, and LLM reranking, usable from the command line or as an MCP server. A large wiki therefore ends up relying on RAG components, applied to already-summarized pages rather than raw documents. The two approaches complement each other more than they exclude each other.

In practice, keep a traditional RAG when the corpus is large or constantly changing (enterprise documentation, tickets, contracts), when the answer must reproduce the exact passage from a document, or when multiple people with different permissions query the same database. LLM Wiki is better suited to a topic you explore for weeks: monitoring, research, reading a book, or preparing a case file. These are uses that the gist itself cites.

#Limits to know, especially locally

Errors accumulate too
A summary error written to a page will be reread, cited, and propagated to subsequent pages. In a RAG, a bad response disappears with the conversation; in a wiki, it remains. That's why systematically linking back to raw sources and reviewing changes are essential.
A local model has less headroom
Ingestion requires following a long instruction, reading several files, and modifying about ten of them without missing any. Small models generally handle this kind of long task less reliably than the large models hosted behind the agents cited in the gist. Judge based on your own sources, starting small.
The context window limits everything
A very long source, a large index, and ten pages to review do not always fit within 64,000 tokens. Split large sources by chapter and keep pages short.
Ingestion takes time
Each source triggers a series of reads and writes. On a modest machine, expect to process sources one at a time rather than import d'ing an entire library in one evening.
The structure drifts
Without strict rules, the agent creates duplicates (the same entity under two names) and pages that nothing links to. The naming rules in the conventions file and the verification pass exist for this purpose.
!
The wiki is not the source of truth
The gist is explicit: the raw sources are authoritative. A wiki page is a summary written by a model. Before relying on a figure, date, or quotation, trace it back to the raw/ file listed as the reference.
!
Saved web pages and hidden instructions
An agent that reads a saved article also reads any instructions contained in that text, and it has permission to write to your files. According to the OpenCode documentation, most actions are authorized by default without confirmation; the rule "permission": { "*": "ask" } in opencode.json requires validation before every action. Keep the wiki in a dedicated folder under git, and review changes after each ingestion of an external source.

#Tips and troubleshooting

The agent skips ingestion steps
The context is probably too short: the instruction falls out of the active window partway through. Check the value of OLLAMA_CONTEXT_LENGTH, shorten AGENTS.md, or split the source.
The agent describes what it would do without writing anything
The model handles tool calls poorly. Choose a model that lists tool support in the Ollama library.
The same entity appears under two names
Ask for a targeted verification pass for duplicates, approve the merges one by one, then add the naming rule that was missing from the conventions file.
The index becomes too long
Split it by category, with a primary index that points to secondary indexes, or add a search tool for Markdown files, such as qmd, which the gist mentions.
The answers ignore existing pages
The index was not updated during ingestion. Rebuild it from the contents of the wiki/ directory, then check the log.

#Ready-to-use implementations

You don't have to write everything by hand. Hermes Agent, Nous Research's open-source agent, documents an integrated skill named llm-wiki, listed under its research category, that follows this pattern. If you're already using this agent with Ollama, it's a faster starting point; the documentation page linked below explains how it works. The principle remains the same: read the conventions before entrusting your sources to them.


#Sources

Everything said about the framework comes from Andrej Karpathy's post and gist. The Ollama and OpenCode settings come from their documentation, already cited in our “OpenCode + Ollama” guide. Review these pages before pasting a command: these tools evolve quickly.

Andrej Karpathy: post on X (April 2, 2026)
https://x.com/karpathy/status/2039805659525644595
Andrej Karpathy: “LLM Wiki” gist (April 4, 2026)
https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
Hermes Agent: llm-wiki skill
https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-llm-wiki
Ollama: context length
https://docs.ollama.com/context-length
Ollama: OpenCode integration
https://docs.ollama.com/integrations/opencode

#Go further

The LLM Wiki sits at the intersection of several topics already covered on the site. Each of these guides covers what this one intentionally leaves out.

Local RAG: introduction
Embeddings, vector database, chunking: how classic RAG works, so you know when it remains the right choice. https://quelllm.fr/guide/rag-local-introduction
Obsidian + local LLM
Connect a local model to an Obsidian vault with the Copilot and Smart Connections plugins, so you can chat with notes you write yourself. https://quelllm.fr/guide/obsidian-llm-local-ollama
NotebookLM locally
Open-source tools that reproduce source notebooks and cited answers, without an intermediate wiki. https://quelllm.fr/guide/notebooklm-local-alternative
Fine-tuning vs. RAG
In his post, Karpathy mentions fine-tuning a model on data from its database as an avenue for exploration. This guide helps you decide whether it is worth the effort. https://quelllm.fr/guide/fine-tuning-vs-rag-choisir
OpenCode + Ollama
The installation of the agent used here, context settings, and permissions. https://quelllm.fr/guide/opencode-ollama-agent-terminal
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.