Karpathy’s LLM Wiki: a knowledge base locale
LLM Wiki is a pattern described by Andrej Karpathy in April 2026: instead of searching through your documents for every question, a model writes and maintains a Markdown wiki that grows with each new source. This guide explains the principle based on his post and gist, proposes a local setup with Ollama and a terminal agent, then clarifies what this pattern does not replace in RAG. It contains neither an in-house test nor a numerical comparison: statements about the pattern are attributed to their source, while the rest is identified as our implementation.
#LLM Wiki: the principle in two minutes
On April 2, 2026, Andrej Karpathy described on X how he uses language models to build personal knowledge bases. Two days later, he published a gist titled “LLM Wiki,” which he presented as a pattern for building this type of database with an LLM. Both texts are short and take ten minutes to read; the links appear at the end of the page.
The gist starts with an observation: most use cases combining LLMs and documents look like RAG. You upload files, the system retrieves excerpts when you ask a question, and the model writes a response. Karpathy acknowledges that this works, but points out that the model rediscovers the knowledge with every question and nothing accumulates. A question requiring five documents to be cross-referenced means finding and piecing together the same parts every time.
The LLM Wiki moves the work upstream. When a new source arrives, the model doesn't just index it: it reads it, extracts the essentials, and integrates them into a set of interconnected Markdown pages. It updates existing pages, revises summaries, and notes where the new source contradicts what was written. The gist describes a persistent artifact that improves over time: the cross-references are already done when the question arrives.
The division of roles is clear in the text. The human chooses the sources, explores, and asks the questions. The model does everything else: summarizing, connecting, classifying, and maintaining the records. Karpathy says he works with the open agent on one side of the screen and Obsidian on the other, and sums up the setup with an image: Obsidian is the IDE, the LLM is the programmer, and the wiki is the codebase.
#Three layers: sources, wiki, conventions
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
The gist describes a three-layer architecture. None requires any particular tool: they are directories and text files.
- The raw sources
- Your document collection: articles, research papers, images, and data files. They are immutable: the model reads them and never modifies them. They—not the wiki—are authoritative.
- The wiki
- A folder of Markdown files written by the model: source summaries, entity pages, concept pages, comparisons, and an overall synthesis. This layer belongs to the model, which creates the pages, updates them, and maintains the links. You read; it writes.
- The conventions file
- A document that tells the model how the wiki is structured, which rules to follow, and how to ingest a source, answer a question, or clean things up. The gist cites CLAUDE.md for Claude Code and AGENTS.md for Codex. This file is what turns a general-purpose agent into a disciplined wiki maintainer.
Two special files help the model—and you—find your way around. The first, index.md, is a catalog: every page appears there with a link and a one-line summary, organized by category. To answer a question, the model first reads the index, then opens the relevant pages. The second, log.md, is a chronological, append-only log: ingestions, questions, and verification passes.
The gist suggests starting each log entry with a consistent prefix, which makes it possible to filter the log with simple Unix tools. The example has this form:
#Three operations: ingest, query, verify
- Ingest
- You drop a source into the raw sources folder and ask the model to process it. According to the gist, it reads the source, discusses the key points with you, writes a summary page, updates the index as well as the relevant entity and concept pages, then adds an entry to the log. Karpathy says a single source can affect 10 to 15 wiki pages.
- Query
- You ask a question. The model searches for the relevant pages, reads them, and writes an answer that cites its references. The gist emphasizes one point: a good answer can be filed in the wiki as a new page, so your explorations accumulate instead of disappearing into a discussion’s history.
- Check (lint)
- From time to time, you ask the model to perform a wiki health check: contradictions between pages, claims superseded by newer sources, orphaned pages with no inbound links, concepts mentioned without a dedicated page, and missing cross-references.
Why entrust this work to a model? The gist's argument is simple: what kills personal wikis is neither reading nor thinking, but record-keeping. Updating cross-references, keeping summaries current, flagging contradictions: the maintenance burden grows faster than the wiki's value, and people give up. A model doesn't get tired and can modify fifteen files in one pass. Karpathy links the idea to the Memex envisioned by Vannevar Bush in 1945.
#Requirements for a wiki maintained by a local model
The gist assumes no specific provider. It needs an agent capable of reading and writing files, driven by a model. Locally, this yields the following building blocks.
- Ollama, up to date
- It serves the model on http://localhost:11434. The ollama launch command used below exists only in recent versions. Installation is covered in our “Install Ollama” guide.
- An agent with file access
- A chat interface is not enough: you need a tool that opens, creates, and modifies files on disk. The example below uses OpenCode, an open-source terminal agent mentioned in the gist, which reads an AGENTS.md file placed at the root of the directory.
- A model that can call tools
- Reading and writing files happens through tool calls. Choose a model that lists tools capability in the Ollama library. Our “OpenCode + Ollama” guide lists several, including qwen3-coder:30b, devstral-small-2:24b, and gpt-oss:20b.
- Memory for context
- Q4_K_M reference figures for weights alone: about 5 GB for 7 billion parameters, 9 GB for 14 billion, and 19 GB for 32 billion. The Ollama documentation requires at least 64,000 tokens of context for agents, which must be added to these figures.
- Git
- The wiki is just a folder of Markdown files: the gist points out that turning it into a Git repository provides version history without adding anything.
- Obsidian (optional)
- To read the wiki, follow links, and display the page graph. Any Markdown editor will do; Obsidian is not involved in writing.
#Step-by-step setup with Ollama
The commands are written for macOS and Linux; on Windows, the simplest option is to use WSL. The directory tree and conventions file are examples to adapt: the gist specifies that the folder structure, conventions, and page format depend on your domain and model, and that everything there is optional and modular.
- 01Create the folder and repositoryOne folder for raw sources, one folder for pages, an index, and a log, all under git.
- 02Write the conventions fileAn AGENTS.md at the root describing the structure, writing rules, and the three procedures: ingestion, querying, and verification.
- 03Launch the model and agentOllama serves a tool-capable model with 64,000 tokens of context; OpenCode opens in the wiki folder.
- 04Ingest a first sourceA single document, processed before your eyes, reviewed, then committed to git.
- 05Query and organize the answersAsk questions in the wiki; useful summaries become pages.
- 06Check regularlyA review pass that lists contradictions, orphaned pages, and broken links.
#1. Create the folder and repository
The raw/ folder will receive your documents, and the wiki/ folder the pages written by the model. Names are up to you, as long as the separation between the two remains clear.
#2. Write the conventions file
This is the part that matters most. Without it, the agent improvises a different structure in every session. Create an AGENTS.md file at the root of ~/wiki, for example based on this:
This file is ours, not Karpathy's: the gist describes the role of the conventions file without providing a template, and recommends evolving it with the model as you learn what works in your domain. Keep it short. The agent rereads it every session, and every line takes up context space.
#3. Launch the model and agent
Download a model capable of calling tools, then open OpenCode in the wiki directory. The command ollama launch opencode starts OpenCode with a model served by Ollama, selected from the picker. OpenCode installation itself is described in our “OpenCode + Ollama” guide.
The remaining issue is context. According to the Ollama documentation, the default window depends on VRAM (4,000 tokens below 24 GB, 32,000 between 24 and 48 GB), while agents require at least 64,000. The OLLAMA_CONTEXT_LENGTH variable sets it when the server starts; if Ollama is already running as an application or service, set the value in its settings instead of starting a second server.
The ollama ps command indicates whether the model fits entirely on the GPU. If it spills over to the processor, every ingestion becomes very slow: choose a smaller model before cutting back on context.
#4. Ingest a first source
Drop an initial document into raw/, preferably in Markdown or plain text. For web pages, the gist points to the Obsidian Web Clipper extension, which converts an article into a Markdown file. Then give the agent the instruction:
The agent reads the source, proposes its key points, then creates and edits the pages. Review the result before going any further: the summary page, the created entity pages, the index, and the log. Then save the wiki state.
#5. Query and organize the responses
After adding a few sources, ask the agent your questions in the same folder. Explicitly ask it to cite its pages and sources, and to say what the wiki does not contain.
#6. Check regularly
Every few ingestion runs, perform a verification pass. Ask for a list of issues before making any corrections: you stay in control of what is merged, renamed, or deleted.
You can inspect the log without an agent. With the regular entry prefix, the command provided in the gist displays the latest operations (only the path is adapted to our directory structure):
#LLM Wiki or RAG: what the prompt cannot replace
The gist contrasts a wiki with RAG to explain the idea. It does not say that one replaces the other, and this guide does not say that either: the two approaches address different situations. Here is what separates them, without numbers, because we have no measurements to present.
- The work moment
- A RAG works question by question: it searches for excerpts, then the model writes the response. A wiki works at ingestion time: the summary is written once, then reread for each question.
- What is retained
- A RAG stores excerpts and their vectors, unreadable as-is. A wiki stores written pages that you can read, edit, and version.
- L'infrastructure
- A RAG system requires an embeddings model, a vector database, and a chunking strategy. The wiki requires a folder and an agent. According to the gist, the index is sufficient at a moderate scale (on the order of a hundred sources and a few hundred pages) and avoids setting up an embeddings-based RAG infrastructure.
- Source fidelity
- A RAG gives the model original passages. The wiki gives it a reformulation written by a model, with the risk of error that entails.
- Volume
- A RAG is designed for large corpora. The wiki is bounded by what the model can read at once: the index, relevant pages, and source must fit within the context window.
Beyond a moderate scale, the gist brings search back in. It cites qmd, a local search engine for Markdown files that combines BM25, vector search, and LLM reranking, usable from the command line or as an MCP server. A large wiki therefore ends up relying on RAG components, applied to already-summarized pages rather than raw documents. The two approaches complement each other more than they exclude each other.
In practice, keep a traditional RAG when the corpus is large or constantly changing (enterprise documentation, tickets, contracts), when the answer must reproduce the exact passage from a document, or when multiple people with different permissions query the same database. LLM Wiki is better suited to a topic you explore for weeks: monitoring, research, reading a book, or preparing a case file. These are uses that the gist itself cites.
#Limits to know, especially locally
- Errors accumulate too
- A summary error written to a page will be reread, cited, and propagated to subsequent pages. In a RAG, a bad response disappears with the conversation; in a wiki, it remains. That's why systematically linking back to raw sources and reviewing changes are essential.
- A local model has less headroom
- Ingestion requires following a long instruction, reading several files, and modifying about ten of them without missing any. Small models generally handle this kind of long task less reliably than the large models hosted behind the agents cited in the gist. Judge based on your own sources, starting small.
- The context window limits everything
- A very long source, a large index, and ten pages to review do not always fit within 64,000 tokens. Split large sources by chapter and keep pages short.
- Ingestion takes time
- Each source triggers a series of reads and writes. On a modest machine, expect to process sources one at a time rather than import d'ing an entire library in one evening.
- The structure drifts
- Without strict rules, the agent creates duplicates (the same entity under two names) and pages that nothing links to. The naming rules in the conventions file and the verification pass exist for this purpose.
#Tips and troubleshooting
- The agent skips ingestion steps
- The context is probably too short: the instruction falls out of the active window partway through. Check the value of OLLAMA_CONTEXT_LENGTH, shorten AGENTS.md, or split the source.
- The agent describes what it would do without writing anything
- The model handles tool calls poorly. Choose a model that lists tool support in the Ollama library.
- The same entity appears under two names
- Ask for a targeted verification pass for duplicates, approve the merges one by one, then add the naming rule that was missing from the conventions file.
- The index becomes too long
- Split it by category, with a primary index that points to secondary indexes, or add a search tool for Markdown files, such as qmd, which the gist mentions.
- The answers ignore existing pages
- The index was not updated during ingestion. Rebuild it from the contents of the wiki/ directory, then check the log.
#Ready-to-use implementations
You don't have to write everything by hand. Hermes Agent, Nous Research's open-source agent, documents an integrated skill named llm-wiki, listed under its research category, that follows this pattern. If you're already using this agent with Ollama, it's a faster starting point; the documentation page linked below explains how it works. The principle remains the same: read the conventions before entrusting your sources to them.
#Sources
Everything said about the framework comes from Andrej Karpathy's post and gist. The Ollama and OpenCode settings come from their documentation, already cited in our “OpenCode + Ollama” guide. Review these pages before pasting a command: these tools evolve quickly.
#Go further
The LLM Wiki sits at the intersection of several topics already covered on the site. Each of these guides covers what this one intentionally leaves out.
- Local RAG: introduction
- Embeddings, vector database, chunking: how classic RAG works, so you know when it remains the right choice. https://quelllm.fr/guide/rag-local-introduction
- Obsidian + local LLM
- Connect a local model to an Obsidian vault with the Copilot and Smart Connections plugins, so you can chat with notes you write yourself. https://quelllm.fr/guide/obsidian-llm-local-ollama
- NotebookLM locally
- Open-source tools that reproduce source notebooks and cited answers, without an intermediate wiki. https://quelllm.fr/guide/notebooklm-local-alternative
- Fine-tuning vs. RAG
- In his post, Karpathy mentions fine-tuning a model on data from its database as an avenue for exploration. This guide helps you decide whether it is worth the effort. https://quelllm.fr/guide/fine-tuning-vs-rag-choisir
- OpenCode + Ollama
- The installation of the agent used here, context settings, and permissions. https://quelllm.fr/guide/opencode-ollama-agent-terminal
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.