Intermediate 11 minWeb search

Perplexica: a self-hosted Perplexity with your LLM local

Perplexity popularized a use case: ask a question in natural language and receive a concise answer with cited sources instead of a list of blue links. Perplexica reproduces exactly this behavior, but on your own machine: an open-source web search engine that queries the web through SearXNG, then has your local LLM under Ollama write the answer. This guide builds the complete stack in Docker, explains the search modes (the famous “focus” modes), and lays out plainly what you gain and lose compared with the original.

By Thomas P.·Update 2026-08-27·Tested on Windows, macOS, and Linux
i
In brief
Perplexica runs Perplexity locally: it queries the web through SearXNG, then has your LLM Ollama write the answer with cited sources. · The stack consists of three components (SearXNG, Ollama, Perplexica) and installs in Docker in about ten minutes. · You need two Ollama models: a chat model (7-14B) for writing, and an embedding model such as nomic-embed-text for ranking sources. · Open-source project (MIT license); no request leaves your network.

#Why Perplexica instead of Perplexity

Perplexity is an excellent product, but it is a cloud service: your queries go to them, the free version is limited, and you do not control the model that writes the response. Perplexica reverses the logic. It is an open-source project (MIT license, available on GitHub) that orchestrates three components you host yourself. No query leaves your network, there is nothing to pay, and you choose the model that responds.

The point isn’t only ideological. A local answer engine can query sources without being constrained by quotas, run in the background in a homelab, and serve as a research component for other tools through its interface. The tradeoff is that quality depends directly on your local model and the health of your SearXNG instance—two points this guide emphasizes.

i
What Perplexica actually does
You ask a question. Perplexica reformulates your query, runs it through SearXNG, retrieves the best results, extracts the relevant content, and then asks your LLM to write a summary citing the numbered sources. This is a RAG pipeline applied to the web in real time.

#The Perplexica + SearXNG + Ollama stack

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Three components, each with a clear role. Understanding what each one does makes troubleshooting much simpler when a response is wrong: you’ll know which of the three to question.

SearXNG
An open metasearch engine that aggregates results from dozens of engines (Google, Bing, DuckDuckGo, Wikipedia…) without tracking. It is what actually searches the web. Perplexica does not work without it.
Ollama
The daemon that serves your local LLM on http://localhost:11434. It produces the final answer from the sources. It also provides an embeddings model for ranking the relevance of excerpts.
Perplexica
The orchestrator and web interface. It reformulates the question, drives SearXNG, extracts page content, handles embedding-based reranking, and composes the final prompt sent to Ollama.
→
Two models, not one
Perplexica needs two separate models: a chat model (which writes) and an embeddings model (which measures similarity to sort the excerpts). Plan to download both into Ollama before configuring the application.

#Prerequisites

Nothing exotic, but a few hardware and version points to watch.

Docker + Docker Compose
The entire stack runs in containers. Docker Desktop (Windows/Mac) or Docker Engine plus the compose plugin (Linux) are sufficient. This is the officially recommended installation method.
Ollama installed and running
The daemon must be running and reachable. Check with ollama list. If Ollama is running on the host machine and Perplexica is running in Docker, the address to use will be http://host.docker.internal:11434, not localhost.
A chat model
A 7–8B in Q4_K_M (~5 GB of VRAM) is a solid baseline; a 14B (~9 GB) produces significantly more reliable summaries. The quality of the final answer depends directly on this choice.
An embeddings model
Lightweight and fast, such as nomic-embed-text. It is used only for internal reranking, not generation.
~5 GB of disk space and a little RAM
For the Docker images (Perplexica + SearXNG) and the cache. The models themselves are managed separately by Ollama.
Terminal
# Récupérer les deux modèles nécessaires côté Ollama
ollama pull qwen3.5:9b         # modèle de chat (rédaction)
ollama pull nomic-embed-text   # modèle d'embeddings (reclassement)

# Vérifier qu'ils sont bien là et qu'Ollama répond
ollama list

#Docker installation in 10 minutes

The official installation uses the project's Git repository, which includes a preconfigured docker-compose.yaml for both Perplexica and a dedicated SearXNG instance. You do not need to install SearXNG separately: compose handles it.

  1. 01
    Clone the repository
    Retrieve the Perplexica project from GitHub and move into its directory. Everything that follows happens there.
  2. 02
    Create the configuration file
    The repository provides a sample.config.toml file. Copy it to config.toml: this is the file Perplexica reads at startup to identify your backends.
  3. 03
    Check the Compose File
    The docker-compose.yaml defines three services: perplexica (the app), searxng, and the required networking. By default, the web interface is exposed on port 3000.
  4. 04
    Launch the stack
    A docker compose up builds the images on the first launch, then starts everything. Allow a few minutes the first time while the images download and build.
  5. 05
    Open the interface
    Go to http://localhost:3000. If the chat page appears, half the work is done; you just need to connect the models.
Terminal
# 1. Cloner le projet officiel
git clone https://github.com/ItzCrazyKns/Perplexica.git
cd Perplexica

# 2. Préparer la config à partir du modèle fourni
cp sample.config.toml config.toml

# 3. Construire et démarrer toute la stack (Perplexica + SearXNG)
docker compose up -d

# 4. Suivre les logs si besoin
docker compose logs -f perplexica
!
Check the exact name of the compose file
Depending on the repository version, the file may be named docker-compose.yaml or docker-compose.yml, and the exposed service may have a slightly different name. Follow the README in the cloned repository, not a memorized command: the project is evolving quickly.

#Connect Ollama to Perplexica

Two ways to configure models: through the config.toml file before launch, or through the web interface's settings screen afterward. The interface is simpler for getting started; the file is useful for reproducible deployment.

The sensitive point is the address of Ollama. Perplexica runs in a container; to it, localhost refers to the container itself, not your machine. If Ollama is installed on the host, use host.docker.internal (Windows/Mac, and Linux with the proper host option) instead of localhost.

config.toml
# Extrait de configuration côté Ollama
# Depuis un conteneur, on ne peut PAS utiliser localhost pour joindre l'hôte
[MODELS.OLLAMA]
API_URL = "http://host.docker.internal:11434"

# Sur une machine Linux, si host.docker.internal ne résout pas,
# ajoutez dans docker-compose : extra_hosts: ["host.docker.internal:host-gateway"]
  1. 01
    Open settings
    In the web interface, the settings icon provides access to the model-provider selection. Select Ollama as the provider.
  2. 02
    Enter Ollama's URL
    Enter the API address (host.docker.internal:11434 from Docker). Perplexica then queries Ollama and automatically lists your available models.
  3. 03
    Choose the chat model
    Select the model that will write the responses (e.g., qwen3.5:9b). This is the number 1 quality lever.
  4. 04
    Choosing an embedding model
    Select nomic-embed-text (or an equivalent) to rerank the excerpts. Without valid embeddings, source ranking suffers.
  5. 05
    Ask a first question
    Test it with a recent factual question. If numbered sources appear below the answer, the full pipeline is working.
→
Ollama and Perplexica in the same compose
You can also add Ollama as a service in docker-compose to keep everything together. In that case, the URL becomes http://ollama:11434 (the service name), and you no longer need host.docker.internal.

#Focus modes: academic, videos, Reddit…

This is Perplexica's signature feature. Instead of searching everywhere in the same way, you choose a “focus mode” that directs SearXNG toward a type of source and adapts how it writes. The right mode radically changes relevance.

All Mode
The default mode: general-purpose search across the entire web. Use it for open-ended questions without a specific domain.
Academic Search
Directs research toward scientific sources and research papers. Ideal for a literature review or a highly technical question.
Writing Assistant
A mode without web search: Perplexica uses only the LLM to draft or rephrase. Useful when you do not need external sources.
YouTube Search
Targets videos: the answer draws on relevant YouTube content, with links. Useful for tutorials and demonstrations.
Wolfram Alpha Search
For computational, scientific, and factual-data questions: the answer relies on Wolfram Alpha.
Reddit Search
Searches Reddit discussions: opinions, user experiences, and community debates. Valuable for opinion-based or product questions.
i
The focus applies to SearXNG, not the model
Changing modes mainly changes which SearXNG sources are queried and how Perplexica filters them. The LLM itself does not change: it is still your Ollama model doing the writing. A poor mode produces poor sources, and no brilliant synthesis can make up for irrelevant sources.

#The actual quality of the responses

Let's be honest: with a small local model, summarization is less fluid and less nuanced than with the leading models that power Perplexity. But the pipeline matters just as much as the model. Three levers determine the final output.

The chat model size
This is the dominant factor. A 7–8B model writes correct answers but is sometimes superficial; a 14B model follows citation instructions better and produces cleaner summaries. Beyond that, a 32B model (~19 GB of VRAM) approaches a convincing experience.
SearXNG health
If engines are blocked or rate-limited, SearXNG returns few results, and the response becomes that much poorer. A hollow response is often caused by a lack of sources, not the LLM.
Available context
Perplexica injects excerpts from several pages into the prompt. A short-context model truncates the sources and loses information. Prefer a model with a comfortable context window (8k minimum, more if possible).
→
Diagnosing a disappointing response
Look at the cited sources first. Few or no sources? The problem is SearXNG (blocked search engines, poorly reformulated query). Relevant sources but a weak synthesis? The problem is the model: move up in size or switch models. This reflex keeps you from optimizing the wrong link in the chain.

#Compared with Perplexity: what you lose, what you gain

Perplexica isn’t a perfect clone, and claiming otherwise would be dishonest. Here’s the realistic balance for deciding whether it’s worth it in your case.

What you gain—privacy
No requests leave your network. Neither your research topics nor the answers pass through a third party. For sensitive monitoring, that is decisive.
You save — cost and quotas
Nothing to pay, no query limits. You can run as many searches as your hardware can handle and integrate it with other tools without an API bill.
You gain—control
You choose the model and the engines queried through SearXNG, and you can self-host everything permanently in a homelab.
You lose—the writing finesse
Perplexity's cloud models are larger and more capable. With modest hardware, local synthesis is duller and makes more approximations.
What you lose—the turnkey reliability
Perplexity manages the infrastructure for you. At home, a rate-limited SearXNG or a Ollama that saturates the VRAM is yours to diagnose and fix.
Some functions are lost
No polished mobile app, no equally mature multi-step Pro Search, no proprietary integrations. Perplexica covers the essentials, not all the convenience of the commercial product.

#Troubleshooting common issues

Perplexica can’t see Ollama
Almost always the address. From Docker, replace localhost with host.docker.internal (and on Linux, add extra_hosts with host-gateway). Make sure Ollama is listening properly and accept connections from outside localhost if necessary.
No sources in the responses
SearXNG returns nothing. Open the SearXNG interface directly to test a search: if it fails, some engines are blocked or rate-limited. Reduce the number of active engines or wait.
The model list is empty
Perplexica queries Ollama when loading the settings. If the list is empty, the API URL is wrong or Ollama has no models. Check with ollama list on the host.
Very slow responses
The chat model is probably too large for your VRAM and is spilling over to the CPU. Switch to a lighter quantization (Q4_K_M) or a smaller model; check that the GPU is actually being used.
Embedding errors
The embedding model has not been selected or downloaded. Run a ollama pull nomic-embed-text and select it explicitly in the settings.
Port 3000 is already in use
Another service is using the port. Change the port mapping in docker-compose (e.g., 3001:3000) and reload the stack.

#Go further

Perplexica is only as good as the model and infrastructure that support it. These guides strengthen the foundations it rests on.

What is Ollama, and how does it work?
The daemon that serves your chat model and embeddings on port 11434—the component on which all writing quality depends.
Choose your quantization
Q4_K_M, Q5_K_M, Q8_0: understand the VRAM/quality tradeoff to choose the right synthesis model for your graphics card.
AnythingLLM: production-ready RAG locally
To go beyond web search and build a RAG over your own documents, using the same backend Ollama.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.