Intermediate 12 minInterfaces

NotebookLM locally: open-source alternatives self-hosted

NotebookLM popularized a simple idea: drop in a bundle of documents, ask questions whose answers cite their sources, and generate an audio summary you can listen to like a podcast. The catch is that everything is sent to Google's servers. This guide shows how to get a local NotebookLM with open-source tools connected to a self-hosted LLM: source notebooks, cited answers, and speech synthesis, without a single file leaving your machine.

By Léa B.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#What NotebookLM does, and what can be replicated

Before rebuilding it, we need to understand what NotebookLM really provides. It is not a general-purpose chatbot: it is an “grounded” assistant based on a document corpus you choose. It answers only from those sources, cites its passages, and theoretically refuses to invent anything not found there. The experience consists of three building blocks.

The source notebook
We import PDFs, web pages, notes, or transcripts. They form the single scope from which the assistant is allowed to draw.
The cited responses
Each claim links to the supporting source passage, so you can verify it with one click instead of trusting blindly.
The audio summary (Audio Overview)
Two synthetic voices discuss your documents in podcast style, so you can listen to a summary while walking instead of reading it.

These three functions can now be replicated with open-source components. The first and second are simply well-designed RAG (Retrieval-Augmented Generation): indexing sources, searching for relevant passages, and generating an answer that points to them. The third is a text → dialogue → text-to-speech (TTS) pipeline. Nothing exotic—the whole thing runs on a machine with an entry-level GPU.

i
What we do not replicate exactly
NotebookLM's polished user experience and the voice quality of its Audio Overview remain difficult to match pixel for pixel. The goal here is not to clone the interface, but to recreate the same use cases — sources, citations, and audio — while keeping control of your data.

#Why want a local NotebookLM

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The number-one reason is privacy. A notebook often contains what is most sensitive: unpublished research notes, internal company documents, client files, contracts, medical records, and reports. Uploading them to Google means entrusting a copy to a third party, under terms of use that evolve and over which you have no control. Running NotebookLM locally solves the problem at its root: the files stay on your drive, and inference runs on your hardware.

Data sovereignty
No document or query passes through a third-party service. Essential for GDPR, professional confidentiality, or confidential R&D.
No quota or subscription
Once the stack is in place, you can process as many sources as your disk and patience allow, with no monthly limit.
Offline
The entire pipeline works without an internet connection, which matters when traveling or on an isolated network.
Model control
You choose the LLM, preferred language, quantization, and settings—instead of being stuck with a black-box model.

#Credible open-source alternatives

Several projects explicitly aim to be a local NotebookLM. None is perfect, but they all connect to Ollama and are improving quickly. Here are the ones that hold up in 2026, from closest to NotebookLM to most “hackable.”

Open Notebook
The closest match to the spirit of NotebookLM: notebooks, source management, cited chat, and built-in podcast generation. Open source, containerized, and compatible with Ollama to remain 100% local.
SurfSense
An open-source research assistant focused on multiple sources (documents, web, connectors). Good for aggregating and querying, with support for local LLMs.
Open WebUI / AnythingLLM
Not NotebookLM clones, but two mature RAG interfaces that cover the essentials: document import, chat with citations, and local embedding models. The simplest path for the “sources + Q&A” part.
Podcastfy
A component dedicated solely to audio: turning documents or URLs into a two-voice conversation. It is driven by a local LLM and a TTS engine of your choice.
→
Two strategies
Either choose an “all-in-one” tool that directly targets NotebookLM (Open Notebook), or assemble an RAG interface yourself (Open WebUI) for sources and citations, then use a separate TTS pipeline for audio. The second approach is more modular and reuses components you may already have.

#Prerequisites

Ollama
The daemon that serves models, on http://localhost:11434 by default. See the installation guide if you haven't done so yet.
A generative model
Qwen 3.5 9B (≈6.6 GB of VRAM in Q4_K_M, 256k context, and multimodal) or Mistral Small 24B (≈14 GB) for excellent French; a Granite 4.2 8B (≈5.3 GB) is sufficient on a small setup.
An embeddings model
nomic-embed-text or mxbai-embed-large, available through ollama pull. They're lightweight and run even without a GPU.
Docker
Most alternatives (Open Notebook, Open WebUI, AnythingLLM) deploy in a single container, avoiding dependency conflicts.
A local TTS engine
Piper for audio: fast, lightweight, with French voices. A GPU is not even required for speech synthesis.
Retrieve the models
# Modèle de génération (adaptez à votre VRAM)
ollama pull qwen3.5:9b

# Modèle d'embeddings pour le RAG
ollama pull nomic-embed-text
i
VRAM guidelines in Q4_K_M
Plan for about 2 GB for a 3B, 5 GB for a 7B, 9 GB for a 14B, and 19 GB for a 32B. A RTX 3060 12 GB runs a 14B comfortably; on a Mac with unified memory (M4 Pro 24–48 GB), a 32B remains feasible.

#Step 1: build your source notebook with Ollama

The first building block of a local NotebookLM is the notebook itself: the place where you drop sources and they are indexed. Here we use Open Notebook, which faithfully reproduces this concept. It runs in a container and is configured to talk to your Ollama rather than a cloud service.

  1. 01
    Retrieve the project
    Clone the Open Notebook repository and go to its directory. It provides a ready-to-use docker-compose file.
  2. 02
    Point to Ollama
    In the configuration (environment file), specify Ollama as the model provider and the URL http://localhost:11434 (or http://host.docker.internal:11434 from the container). Choose your generation model and nomic-embed-text for embeddings.
  3. 03
    Start the stack
    Run docker compose up. The web interface then opens in the browser, with the database and indexing managed for you.
  4. 04
    Create a notebook and import
    Create a notebook, then drag your PDFs into it, or paste URLs or text. Each source is automatically split and vectorized.
Terminal
# Récupérer et lancer Open Notebook
git clone https://github.com/lfnovo/open-notebook.git
cd open-notebook

# Configurer le fournisseur Ollama dans le fichier d'environnement
cp .env.example .env
# éditez .env : OLLAMA_API_BASE=http://host.docker.internal:11434

# Démarrer la stack complète
docker compose up -d
!
host.docker.internal on Linux
From a Docker container, localhost refers to the container, not your machine. On macOS and Windows, host.docker.internal points to the host. On Linux, you may need to add it explicitly (extra_hosts: host.docker.internal:host-gateway in the compose file) or use the Docker gateway IP. Otherwise, Ollama will remain unreachable.

If you'd rather not add another tool, Open WebUI and AnythingLLM work very well as source notebooks: you create a workspace, import the documents into it, and indexing via nomic-embed-text happens automatically. This is the “no-code” path detailed in the site's RAG guides.

#Step 2: cited Q&A

This is the core of a local NotebookLM: ask a question in natural language and get an answer based only on your sources, with the exact supporting passage. Technically, RAG retrieves the excerpts closest to the question, then the LLM drafts the answer from them—not from its general knowledge. Citation quality depends mainly on two things: a strict system prompt and explicitly requesting the source.

In both Open Notebook and Open WebUI, this behavior is already built in: ask the question in the notebook chat, and the answer displays the excerpts used. If you build your own pipeline, the system prompt makes all the difference.

System prompt for cited answers
Tu réponds UNIQUEMENT à partir des extraits fournis ci-dessous.

Règles :
- Si la réponse ne figure pas dans les extraits, dis « Je ne trouve pas cette information dans les sources. »
- N'utilise jamais tes connaissances générales pour compléter.
- Après chaque affirmation, indique la source entre crochets, ex. [source 2].
- Cite mot pour mot le passage clé quand c'est utile.

Réponds en français, de façon concise.

For model settings, two parameters matter. A low temperature (0.2) limits embellishment and keeps the model grounded in the sources. A sufficient context window (num_ctx) lets it ingest the retrieved excerpts without truncating them—otherwise, the model answers based on only part of the passages.

Ollama call with sources
import requests

SYSTEM = open('prompt_systeme.txt').read()

# extraits = passages renvoyés par votre recherche vectorielle
contexte = "\n\n".join(
    f"[source {i+1}] {e}" for i, e in enumerate(extraits)
)

resp = requests.post("http://localhost:11434/api/chat", json={
    "model": "qwen3.5:9b",
    "stream": False,
    "options": {"num_ctx": 8192, "temperature": 0.2},
    "messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": f"{contexte}\n\nQuestion : {question}"},
    ],
})

print(resp.json()["message"]["content"])
i
RAG does not read “everything”
Contrary to a common assumption, the model does not scan all your sources for every question: it receives only the few closest passages. A poorly worded question retrieves the wrong excerpts and produces an off-target answer. If the answer is disappointing, rephrase the question using the document’s exact terms.

#Step 3: Generate a local audio summary (TTS)

This is NotebookLM's signature feature—the Audio Overview, in which two voices discuss your documents—and it's also the most impressive feature to recreate locally. The pipeline has two stages: the LLM writes a dialogue script from the sources, then a local TTS engine turns that script into audio. This keeps both the voice and the text fully offline.

First step: have the conversation written. Ask the LLM for a two-character exchange that explains the content in plain language, marking each line with its speaker — this labeling will be used to alternate voices during synthesis.

Podcast script prompt
À partir des sources ci-dessous, écris un dialogue de podcast en français
entre deux animateurs, Alex et Camille, qui vulgarisent le contenu.

Règles :
- Format strict, une réplique par ligne : « Alex: ... » ou « Camille: ... ».
- Reste fidèle aux sources, n'invente aucun fait ni chiffre.
- Ton vivant et curieux, phrases courtes, adaptées à l'oral.
- 12 à 18 répliques, une vraie conversation qui se répond.

Second step: speech synthesis. Piper is the natural choice for local use: fast, lightweight, and with good-quality French voices. We assign one voice per speaker (two different .onnx model files) to distinguish Alex and Camille, then concatenate the audio segments.

Install Piper and a French voice
# Installer Piper
pip install piper-tts

# Télécharger deux voix françaises depuis Hugging Face (rhasspy/piper-voices)
# ex. fr_FR-siwis-medium et fr_FR-upmc-medium (fichiers .onnx + .onnx.json)

# Synthétiser une réplique
echo "Bonjour et bienvenue dans cet épisode." | \
  piper --model fr_FR-siwis-medium.onnx --output_file alex_01.wav
script_vers_audio.py
import subprocess, re

VOIX = {
    "Alex": "fr_FR-siwis-medium.onnx",
    "Camille": "fr_FR-upmc-medium.onnx",
}

segments = []
for i, ligne in enumerate(open("script.txt")):
    m = re.match(r"(Alex|Camille):\s*(.+)", ligne.strip())
    if not m:
        continue
    locuteur, texte = m.group(1), m.group(2)
    wav = f"seg_{i:03d}.wav"
    subprocess.run(
        ["piper", "--model", VOIX[locuteur], "--output_file", wav],
        input=texte.encode(),
    )
    segments.append(wav)

# Concaténer les segments (ex. avec ffmpeg ou pydub)
print("Segments générés :", len(segments))

To avoid writing this pipeline by hand, Open Notebook includes podcast generation, and Podcastfy does exactly this “documents → audio conversation” workflow with a local LLM and the TTS engine of your choice. The manual pipeline above is still useful for understanding what is happening and maintaining full control over the voices and formatting.

→
More natural voices
Piper prioritizes speed over expressiveness. If the output sounds too robotic, engines such as Kokoro or Coqui XTTS produce more natural voices, at the cost of heavier synthesis and a nearly mandatory GPU. For utility listening, Piper is more than sufficient.

#Troubleshooting

The container does not reach Ollama
From Docker, localhost points to the container. Use host.docker.internal (added via extra_hosts on Linux) or the host's IP, and check that Ollama is listening on 0.0.0.0 if needed.
Uncited or fabricated responses
The system prompt is not strict enough, or the temperature is too high. Enforce “only from the excerpts,” require the [source N] format, and lower the temperature to 0.2.
The PDF comes out empty during indexing
It's a scan with no text layer. Run it through OCR (ocrmypdf entree.pdf sortie.pdf) before importing it.
The response ignores some of the sources
num_ctx is too small: the excerpts were truncated. Increase it if VRAM allows, or reduce the number of retrieved passages.
Piper cannot find the voice
You need both files for each voice: the .onnx and its accompanying .onnx.json. Check the exact path passed to --model.
Choppy audio between turns
Raw concatenation joins WAV files without any breathing room. Insert a short silence between segments (ffmpeg or pydub) for smoother output.

#Go further

This guide brings together building blocks already covered elsewhere on the site. To explore each step in more depth:

Install Ollama: Windows, macOS, and Linux
The starting point for serving your models locally on port 11434, if you haven't set it up yet.
Local RAG with Ollama without coding (Open WebUI, AnythingLLM)
The no-code route for the “sources + cited questions and answers” part of your local NotebookLM.
100% local voice assistant: Whisper + Ollama + Piper
For going further with Piper and local speech synthesis, beyond audio summaries alone.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.