Open Notebook: the AI research notebook self-hosted
Open Notebook (lfnovo/open-notebook, MIT license) is a self-hosted AI-assisted research notebook: you collect PDFs, web pages, or audio, and the tool indexes them and then answers with citations. It connects to 18 providers of your choice, including Ollama locally, and generates multi-voice podcasts (1 to 4) through customizable episode profiles, whereas NotebookLM limits audio to two fixed voices.
Open Notebook takes Google NotebookLM's idea of an AI-assisted research notebook—gather sources, query them, and turn them into notes and podcasts—and implements it as a self-hosted open-source project (lfnovo/open-notebook repository, MIT license) that you connect to the model of your choice. For anyone who tried the online equivalent and stopped at the question “where do my documents go?”, it offers the same promise without the transfer, along with features Google's version doesn't have.
#The idea: an assisted research notebook
A notebook brings together sources: PDFs, web pages, YouTube videos, audio files, and Office documents. The tool indexes them, lets you query them in natural language, and, above all, produces derived artifacts—summaries, cross-source syntheses, notes, and multi-voice podcasts—that remain linked to their sources. The project presents itself as a private, multi-model, 100% local, fully functional alternative to Notebook LM.
The difference from a simple document assistant lies in this notion of a notebook: you don’t ask a corpus an isolated question; you build a dossier on a topic and work in it over time, with multiple notebooks for several research projects in parallel, each with its own sources and history.
#What the local version changes
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
- Sources do not leave the system
- Confidential documents, client files, embargoed materials: processing stays on your machine if you choose local providers (Ollama, LM Studio, or oMLX on Apple Silicon).
- Choosing the model
- 18 providers supported natively — OpenAI, Anthropic, Google, Mistral, Groq, xAI, DeepSeek, Ollama, LM Studio, OpenRouter, and others — through the Esperanto library by the same author. You decide, provider by provider.
- No arbitrary limits
- The number of notebooks and sources depends on your disk, not on a pricing tier; the actual cost is limited to the API usage you choose to connect.
- In return, the operational side
- This is a Docker service (Python/FastAPI backend, Next.js/React interface, SurrealDB database) to install, back up, and update. The usual trade-off of self-hosting.
#What Open Notebook does beyond NotebookLM
The official repository publishes a direct comparison table with Google NotebookLM. On several points, the difference is more than just “it's local”: the number of voices in generated podcasts, API access, and customizable content transformations go beyond what Google's tool offers.
| Criterion | Open Notebook | Google NotebookLM |
|---|---|---|
| AI providers | 18+ to choose from (including Ollama, LM Studio) | Google models only |
| Generated podcast voice | 1 to 4, customizable episode profiles | 2 fixed voices, fixed format |
| API access | Full REST API | No API |
| Deployment | Docker, cloud, or local | Hosted exclusively by Google |
| Source citations | Basic references (the project says it intends to improve them) | Complete, more polished citations |
This last point deserves an honest mention: in its own comparison published on GitHub, the project acknowledges that its citations are currently more basic than NotebookLM's. It is therefore not a superior replacement on every front—it is a compromise that becomes favorable when privacy, model choice, API-based automation, or podcast generation matter more than the fine quality of the citations displayed on screen.
- Podcasts with episode profiles
- Debate, lecture, interview, explanation for beginners or experts: you configure the episode structure, unlike NotebookLM's single “deep dive” format.
- Reasoning model support
- Full support for “thinking” models such as DeepSeek-R1 or Qwen3, useful for summaries that require multiple reasoning steps.
- Content transformations
- Customizable actions for summarizing, extracting insights, or rephrasing, beyond simple automatic summarization.
- Password protection
- Useful for a publicly exposed deployment, in addition to fine-grained control over what is shared with each model.
#Install locally, with or without Ollama
The documented deployment takes three steps: retrieve the official docker-compose.yml, enter the desired API keys in the interface (no configuration file to edit manually), then start the services. The only stated prerequisite is Docker Desktop.
- 01Retrieve the service filecurl -o docker-compose.yml https://raw.githubusercontent.com/lfnovo/open-notebook/main/docker-compose.yml, ou copier le contenu manuellement depuis le dépôt.
- 02Choose the 100% local variant if neededThe repository provides a separate file, examples/docker-compose-ollama.yml, which adds Ollama as the default provider so nothing is sent over the internet.
- 03Start and configuredocker compose up -d, puis ouvrir l'interface locale et renseigner les clés d'API des fournisseurs distants éventuels ; aucune clé n'est nécessaire pour un usage 100 % Ollama.
The official documentation offers two distinct paths depending on your profile: a “Quick Start with OpenAI” to get up and running in five minutes with a remote provider, and a “Run It Fully Local” guide dedicated to Ollama and LM Studio for completely private use from installation onward. The full REST API has its own interactive documentation on local port 5055, useful for connecting the notebook to a script or existing automation instead of doing everything through the web interface.
#Three practical ways to use it
The project organizes work around several independent notebooks (“Multi-Notebook Organization”), each with its own sources, notes, and discussion history. Three use cases come up most often among users documenting their migration from NotebookLM.
- Literature review
- One notebook per research topic, dozens of PDFs added as you read, and a customized content transformation that systematically extracts the method, results, and limitations from each paper.
- Monitoring and meeting notes
- Regularly added transcripts, a contextual discussion for finding “what was decided about X in March,” and a summary podcast to listen to while working on other tasks.
- Product documentation or customer support
- All internal documentation in a dedicated notebook, queried in natural language by the team, with a local provider so internal documents or sensitive procedures do not pass through a third-party service hosted elsewhere.
What distinguishes these uses from a simple chat with a model is that every answer remains grounded in the sources from the active notebook: switching notebooks changes the available context without requiring you to reload or explain anything to the model. Fine-grained control over what is shared with the model—a subset of sources rather than the entire notebook—also helps limit the context sent, and therefore latency and cost, for a large notebook.
#How to actually use it
- 01One notebook per topic, not one notebook for everythingAnswer quality drops when a notebook mixes unrelated subjects: the passages most closely related to a question begin coming from elsewhere.
- 02Add sources progressivelyAnd verify what the tool actually extracted from each one. A scanned PDF without a text layer is indexed empty, regardless of which search engine is used behind it.
- 03Configure a suitable embedding modelIt determines what the search retrieves, much more than the model that writes the response. For French documents, it must be multilingual; Ollama, Google, and Mistral offer embeddings in the project's provider matrix.
- 04Require citations, keeping their limitations in mindA response without a source cannot be verified. The project itself says that citations are still basic: verifying the cited source remains essential before reusing a result as-is.
#The limitations, including those the project itself acknowledges
- Quality comes from the model
- A small local model summarizes in a flat way. The tool organizes the work; it does not replace the ability to synthesize—hence the value of reasoning-model support for complex tasks.
- Citations are still basic
- The project's own comparison says so: less polished than NotebookLM's, with an announced but undated improvement.
- Audio and video sources require transcription
- Guaranteed depending on the provider chosen (Whisper via APIs, Deepgram, or a local engine such as Faster-Whisper upstream), but it is an additional processing step.
- Young and active project
- The repository started in late 2024 and is evolving quickly; check the announced features against the version you install (CHANGELOG, GitHub releases) rather than relying on an article.
- Not every provider does everything
- The matrix published by the project distinguishes four capabilities by provider—LLM, embedding, transcription, and text-to-speech—and few providers cover them all. Anthropic and Groq, for example, do not offer embeddings in Open Notebook: you often need to combine two providers, one for drafting and another for indexing or transcription.
#Open Notebook or a traditional RAG system
| Need | Tool |
|---|---|
| Build a dossier on a topic and work inside it, with a summary podcast | Open Notebook |
| Answer questions about an enterprise document repository | A dedicated RAG application (AnythingLLM, for example) |
| Integrate document search into an application | A RAG pipeline you write yourself, with your own embeddings |
| Chat with a few PDFs without installing anything else | An all-in-one desktop application |
- AnythingLLM: production-ready RAG locally
- The embedding model determines what you retrieve
- Properly prepare PDFs before indexing
- Transcribe audio sources before adding them to the notebook
- Source: official Open Notebook GitHub repository
- Source: official docker-compose.yml file
- Source: Esperanto library (multi-provider layer)
#FAQ
Is Open Notebook free?+
Does it work completely offline?+
How many voices for a generated podcast?+
Do the answers really cite their sources?+
Why do my answers ignore a source's content?+
Do you need a graphics card?+
What's the difference from an RAG I would build myself?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.