SearXNG: giving a model web search local
SearXNG (free software under the AGPL-3.0 license that neither tracks nor profiles its users) is a self-hosted metasearch engine: it relays your queries to other engines and merges their results, with no API key or bill. For a local model, the step that unlocks everything is adding “json” to the list of formats in settings.yml (“html” alone by default), then reconnecting this JSON output to your chat interface or agent.
An open model knows nothing about this week, and nothing in its response warns you: it will talk about last month with the same confidence as a multiplication table memorized by heart. Giving it access to web search fixes that. SearXNG is a self-hosted metasearch engine that provides this service without an API key, per-query charges, or an account—making it a natural building block for a local setup, where you do not want to send every question to a third party.
#Why a local model needs the web
A model's knowledge ends at its training date, and it doesn't know with certainty where that boundary lies. There are two solutions: provide the facts yourself in the prompt, or give it a way to retrieve them at response time. Search is the general form of the latter, and the most scalable as soon as the questions vary.
The obvious route is a commercial search API: an account, a key, and a bill per query. SearXNG is the self-hosted alternative. You run the instance, queries originate from your system, no account exists, and nothing is logged unless you decide otherwise. For an installation whose entire point is keeping data at home, sending every question to a third-party service through the back door would be a strange place to stop.
The project itself clearly presents itself as a metasearch engine where users are neither tracked nor profiled, and it is distributed under the AGPL-3.0 free software license. This is a strong copyleft license: redistributing a modified version, including as an online service, requires republishing the corresponding source code under the same terms. For personal or internal team use, this clause is never triggered: it only concerns redistribution of a modified version to third parties, not simply running the instance at home or in your company.
#What SearXNG does, without magic
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
SearXNG doesn't crawl the web and has no index. It's an intermediary: your query is sent to a configurable set of existing engines, their responses come back, are deduplicated and reranked, and you receive a single list. The engines queried see the instance's IP address, not yours.
This intermediary status also explains the privacy promise emphasized by the project itself: users are neither tracked nor profiled by the instance. No account, no advertising-tracking cookie, and no search history tied to an identity—which clearly distinguishes SearXNG from commercial search engines whose business model relies precisely on profiling queries and clicks over time.
- Quality is inherited
- Disable engines that return noise for your specific use case: this improves the model’s responses more than any prompt tweak, because the model cannot recover from an already polluted starting context.
- The automation feels like abuse
- A human runs a few searches per minute; an agent loop runs dozens. Upstream engines respond with CAPTCHAs and temporary blocks, which appear as empty results.
This strong dependence on upstream engines is the central trade-off of any metasearch engine: SearXNG cannot be better than what it aggregates, and it also inherits their outages and changes. If an engine changes its HTML markup overnight, the corresponding SearXNG module breaks until the maintainers update it. This explains why an instance that has been left unchanged for a long time may see some engines silently stop responding without an explicit error message.
#L'installer
- 01Deploy the containerThe official image is published on Docker Hub (searxng/searxng) and GHCR (ghcr.io/searxng/searxng), both maintained by the project itself. On a single machine, the image plus a Valkey cache service (the open-source fork of Redis) are more than enough to get started.
- 02Keep it on the local networkAn instance exposed to the internet is found within a few days and used as a free relay, causing upstream services to block its access. Bind it to the local network, or put authentication in front of it.
- 03Check from the machine that will call itNot from your browser. A container calling another needs the address visible from Docker, not localhost.
#The setting everyone misses
By default, a freshly installed SearXNG instance serves only HTML: perfect for a human, completely unusable for a program. The official documentation confirms that the search.formats key in settings.yml contains only “html” by default, while csv, json, and rss are available; you must add json and then restart. This is the sole cause of the vast majority of “it doesn’t work” issues encountered when connecting SearXNG to a chat interface.
#Connect it to your model
| Environment | SearXNG’s role |
|---|---|
| Local chat web interface | Built-in search provider: enter the instance URL, and responses gain a source list |
| Agent framework | A research tool the agent can call |
| Document retrieval pipeline | Retrieval step for current-events questions, alongside your document index |
| Automation tool | A simple HTTP call to the JSON endpoint |
The pattern is always the same: the search returns excerpts and addresses, your code or interface keeps a few of them, and they are inserted into the prompt before the model answers the question. The model doesn't browse: it reads what you hand it. That's why excerpt quality and context-window size matter more here than model size.
#The settings that really make a difference
A freshly installed instance enables a large number of engines by default, which seems generous but dilutes quality: more noise, more duplicates, and more latency while waiting for the slowest engine in the batch to return its response. The first real improvement never involves the model prompt; it involves the list of active engines and their actual relevance to the intended use case.
- Turn off irrelevant engines
- For example, a use case focused on technical documentation does not need image or shopping engines enabled by default. Each engine removed means one less source of noise and one fewer upstream request to manage day to day.
- Enable caching
- A shared Valkey database avoids repeating the same search on every call from an agent that asks the same question repeatedly in a loop, saving both response time and valuable quota with upstream engines.
- Set an output budget
- Asking for ten results when the model will actually read only three wastes a lot of context and latency for no concrete benefit. Align the number of results returned with what your prompt really uses behind the scenes, no more and no less.
- Monitor the logs at startup
- An instance recently upgraded to Valkey can fail silently if the hostname still points to an old cache service. A quick look at the container logs prevents you from looking elsewhere for a problem that is right there.
- Install Open WebUI, the interface that consumes these results
- Automate workflows with n8n and Ollama
- Architecture of a local agent that uses tools
- LiteLLM: a unified local and cloud proxy
#Limitations to plan for
- Engine quotas
- Empty results almost always mean an upstream blockage, not a configuration error on SearXNG itself. Use fewer active engines, more Valkey caching, and above all, no agent that searches in a loop without any rate limit.
- Short excerpts
- For in-depth content, you need to retrieve and clean the pages behind the links: that's the job of a dedicated extractor, not a search engine that merely relays results already summarized upstream.
- No intelligent ranking
- SearXNG merges results; it does not understand your question. A vague query returns vague context, which the model dutifully summarizes without ever pointing out that the question itself was poorly phrased.
- Maintenance comes back to you
- Upstream engines change their pages, so modules must keep up. An instance left without updates for a year degrades silently.
- The Valkey migration applies to older instances
- An instance mounted more than a year ago may still reference an old cache service. Check the hostname configured for the rate limiter: if it no longer matches the container actually running today, the anti-bot protection fails silently instead of properly blocking abuse.
#FAQ
Is SearXNG free?+
Why does my instance return HTML instead of JSON?+
Can I use a public instance instead of hosting my own?+
Do you need a GPU?+
Why do the results become empty after a while?+
Does this bring my model up to date?+
What does SearXNG's AGPL-3.0 license mean for me?+
- Source: SearXNG GitHub repository
- Source: official container installation guide
- Source: results-format documentation
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.