Perplexica: a self-hosted Perplexity with your LLM local
Perplexity popularized a use case: ask a question in natural language and receive a concise answer with cited sources instead of a list of blue links. Perplexica reproduces exactly this behavior, but on your own machine: an open-source web search engine that queries the web through SearXNG, then has your local LLM under Ollama write the answer. This guide builds the complete stack in Docker, explains the search modes (the famous “focus” modes), and lays out plainly what you gain and lose compared with the original.
#Why Perplexica instead of Perplexity
Perplexity is an excellent product, but it is a cloud service: your queries go to them, the free version is limited, and you do not control the model that writes the response. Perplexica reverses the logic. It is an open-source project (MIT license, available on GitHub) that orchestrates three components you host yourself. No query leaves your network, there is nothing to pay, and you choose the model that responds.
The point isn’t only ideological. A local answer engine can query sources without being constrained by quotas, run in the background in a homelab, and serve as a research component for other tools through its interface. The tradeoff is that quality depends directly on your local model and the health of your SearXNG instance—two points this guide emphasizes.
#The Perplexica + SearXNG + Ollama stack
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Three components, each with a clear role. Understanding what each one does makes troubleshooting much simpler when a response is wrong: you’ll know which of the three to question.
- SearXNG
- An open metasearch engine that aggregates results from dozens of engines (Google, Bing, DuckDuckGo, Wikipedia…) without tracking. It is what actually searches the web. Perplexica does not work without it.
- Ollama
- The daemon that serves your local LLM on http://localhost:11434. It produces the final answer from the sources. It also provides an embeddings model for ranking the relevance of excerpts.
- Perplexica
- The orchestrator and web interface. It reformulates the question, drives SearXNG, extracts page content, handles embedding-based reranking, and composes the final prompt sent to Ollama.
#Prerequisites
Nothing exotic, but a few hardware and version points to watch.
- Docker + Docker Compose
- The entire stack runs in containers. Docker Desktop (Windows/Mac) or Docker Engine plus the compose plugin (Linux) are sufficient. This is the officially recommended installation method.
- Ollama installed and running
- The daemon must be running and reachable. Check with ollama list. If Ollama is running on the host machine and Perplexica is running in Docker, the address to use will be http://host.docker.internal:11434, not localhost.
- A chat model
- A 7–8B in Q4_K_M (~5 GB of VRAM) is a solid baseline; a 14B (~9 GB) produces significantly more reliable summaries. The quality of the final answer depends directly on this choice.
- An embeddings model
- Lightweight and fast, such as nomic-embed-text. It is used only for internal reranking, not generation.
- ~5 GB of disk space and a little RAM
- For the Docker images (Perplexica + SearXNG) and the cache. The models themselves are managed separately by Ollama.
#Docker installation in 10 minutes
The official installation uses the project's Git repository, which includes a preconfigured docker-compose.yaml for both Perplexica and a dedicated SearXNG instance. You do not need to install SearXNG separately: compose handles it.
- 01Clone the repositoryRetrieve the Perplexica project from GitHub and move into its directory. Everything that follows happens there.
- 02Create the configuration fileThe repository provides a sample.config.toml file. Copy it to config.toml: this is the file Perplexica reads at startup to identify your backends.
- 03Check the Compose FileThe docker-compose.yaml defines three services: perplexica (the app), searxng, and the required networking. By default, the web interface is exposed on port 3000.
- 04Launch the stackA docker compose up builds the images on the first launch, then starts everything. Allow a few minutes the first time while the images download and build.
- 05Open the interfaceGo to http://localhost:3000. If the chat page appears, half the work is done; you just need to connect the models.
#Connect Ollama to Perplexica
Two ways to configure models: through the config.toml file before launch, or through the web interface's settings screen afterward. The interface is simpler for getting started; the file is useful for reproducible deployment.
The sensitive point is the address of Ollama. Perplexica runs in a container; to it, localhost refers to the container itself, not your machine. If Ollama is installed on the host, use host.docker.internal (Windows/Mac, and Linux with the proper host option) instead of localhost.
- 01Open settingsIn the web interface, the settings icon provides access to the model-provider selection. Select Ollama as the provider.
- 02Enter Ollama's URLEnter the API address (host.docker.internal:11434 from Docker). Perplexica then queries Ollama and automatically lists your available models.
- 03Choose the chat modelSelect the model that will write the responses (e.g., qwen3.5:9b). This is the number 1 quality lever.
- 04Choosing an embedding modelSelect nomic-embed-text (or an equivalent) to rerank the excerpts. Without valid embeddings, source ranking suffers.
- 05Ask a first questionTest it with a recent factual question. If numbered sources appear below the answer, the full pipeline is working.
#Focus modes: academic, videos, Reddit…
This is Perplexica's signature feature. Instead of searching everywhere in the same way, you choose a “focus mode” that directs SearXNG toward a type of source and adapts how it writes. The right mode radically changes relevance.
- All Mode
- The default mode: general-purpose search across the entire web. Use it for open-ended questions without a specific domain.
- Academic Search
- Directs research toward scientific sources and research papers. Ideal for a literature review or a highly technical question.
- Writing Assistant
- A mode without web search: Perplexica uses only the LLM to draft or rephrase. Useful when you do not need external sources.
- YouTube Search
- Targets videos: the answer draws on relevant YouTube content, with links. Useful for tutorials and demonstrations.
- Wolfram Alpha Search
- For computational, scientific, and factual-data questions: the answer relies on Wolfram Alpha.
- Reddit Search
- Searches Reddit discussions: opinions, user experiences, and community debates. Valuable for opinion-based or product questions.
#The actual quality of the responses
Let's be honest: with a small local model, summarization is less fluid and less nuanced than with the leading models that power Perplexity. But the pipeline matters just as much as the model. Three levers determine the final output.
- The chat model size
- This is the dominant factor. A 7–8B model writes correct answers but is sometimes superficial; a 14B model follows citation instructions better and produces cleaner summaries. Beyond that, a 32B model (~19 GB of VRAM) approaches a convincing experience.
- SearXNG health
- If engines are blocked or rate-limited, SearXNG returns few results, and the response becomes that much poorer. A hollow response is often caused by a lack of sources, not the LLM.
- Available context
- Perplexica injects excerpts from several pages into the prompt. A short-context model truncates the sources and loses information. Prefer a model with a comfortable context window (8k minimum, more if possible).
#Compared with Perplexity: what you lose, what you gain
Perplexica isn’t a perfect clone, and claiming otherwise would be dishonest. Here’s the realistic balance for deciding whether it’s worth it in your case.
- What you gain—privacy
- No requests leave your network. Neither your research topics nor the answers pass through a third party. For sensitive monitoring, that is decisive.
- You save — cost and quotas
- Nothing to pay, no query limits. You can run as many searches as your hardware can handle and integrate it with other tools without an API bill.
- You gain—control
- You choose the model and the engines queried through SearXNG, and you can self-host everything permanently in a homelab.
- You lose—the writing finesse
- Perplexity's cloud models are larger and more capable. With modest hardware, local synthesis is duller and makes more approximations.
- What you lose—the turnkey reliability
- Perplexity manages the infrastructure for you. At home, a rate-limited SearXNG or a Ollama that saturates the VRAM is yours to diagnose and fix.
- Some functions are lost
- No polished mobile app, no equally mature multi-step Pro Search, no proprietary integrations. Perplexica covers the essentials, not all the convenience of the commercial product.
#Troubleshooting common issues
- Perplexica can’t see Ollama
- Almost always the address. From Docker, replace localhost with host.docker.internal (and on Linux, add extra_hosts with host-gateway). Make sure Ollama is listening properly and accept connections from outside localhost if necessary.
- No sources in the responses
- SearXNG returns nothing. Open the SearXNG interface directly to test a search: if it fails, some engines are blocked or rate-limited. Reduce the number of active engines or wait.
- The model list is empty
- Perplexica queries Ollama when loading the settings. If the list is empty, the API URL is wrong or Ollama has no models. Check with ollama list on the host.
- Very slow responses
- The chat model is probably too large for your VRAM and is spilling over to the CPU. Switch to a lighter quantization (Q4_K_M) or a smaller model; check that the GPU is actually being used.
- Embedding errors
- The embedding model has not been selected or downloaded. Run a ollama pull nomic-embed-text and select it explicitly in the settings.
- Port 3000 is already in use
- Another service is using the port. Change the port mapping in docker-compose (e.g., 3001:3000) and reload the stack.
#Go further
Perplexica is only as good as the model and infrastructure that support it. These guides strengthen the foundations it rests on.
- What is Ollama, and how does it work?
- The daemon that serves your chat model and embeddings on port 11434—the component on which all writing quality depends.
- Choose your quantization
- Q4_K_M, Q5_K_M, Q8_0: understand the VRAM/quality tradeoff to choose the right synthesis model for your graphics card.
- AnythingLLM: production-ready RAG locally
- To go beyond web search and build a RAG over your own documents, using the same backend Ollama.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.