Secure your Ollama server: authentication, reverse proxy, exposition
Securing Ollama is not optional when the daemon becomes accessible beyond your machine. Thousands of instances are left exposed on the internet without any access control, offering free GPU compute to anyone who finds them—and sometimes much worse. This guide shows how to check your exposure, put authentication in front of the API through a reverse proxy, encrypt traffic, and open remote access properly.
#Why securing Ollama is urgent
Ollama has no built-in authentication. The daemon listens, serves requests, and that's it: it assumes that only a trusted client on the local machine is talking to it. This model holds as long as you stay on http://localhost:11434. The problem appears as soon as you change an environment variable to make the service accessible over the network—a common step when connecting a remote interface or another machine.
Setting OLLAMA_HOST to 0.0.0.0 tells the daemon to listen on all network interfaces. If port 11434 is not blocked by a firewall, the API is exposed to the entire local network, or even the entire internet if the machine has a public IP or port forwarding on the router. There is no password and no token: anyone can send requests, download or delete your models, and consume your GPU.
#What the Ollama API really exposes
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
Before protecting anything, you need to understand the scope of the attack surface. The Ollama API is more than a chat endpoint: it is a full administration API with no privilege separation. An anonymous client has exactly the same rights as you.
- /api/generate et /api/chat
- Text generation. Uses your GPU as much as it wants and sees every prompt—including potentially sensitive data injected by your own applications.
- /api/tags
- Lists every installed model. An attacker immediately knows what you are hosting, including any internal fine-tunes you may have.
- /api/pull
- Download any model from a registry. A third party could fill your disk or install a compromised model.
- /api/delete
- Deletes models. Possible data destruction without any authentication.
- /api/create et /api/push
- Create models from a Modelfile and send them to a registry. Full takeover of the instance.
- /v1/*
- OpenAI-compatible layer on the same port. The same lack of access control, exploitable by any standard OpenAI client.
#Check whether your Ollama is exposed
Start with an honest diagnosis. The goal: determine which interfaces the daemon is listening on and whether the port is reachable from outside. Three checks, from the most local to the most external.
First, check which addresses the process is listening on. Listening on 127.0.0.1 is healthy; listening on 0.0.0.0 or a network IP means the service accepts remote connections.
Next, test from another machine on the network. Replace IP_DU_SERVEUR with the local address of the Ollama machine. If the command returns the model list, the instance is reachable on the network—which is acceptable on a trusted LAN, but never without filtering toward the internet.
Finally, check your public exposure. If your machine has a public IP or an active port forwarding rule, look it up on an indexing service such as Shodan or Censys. A simple query on port 11434 and the Ollama banner reveals whether your instance is already indexed. You can also test your public IP directly.
#Step 1 — Switch back to listening on localhost
The first step, and often the only one needed, is to restore Ollama to its default behavior: listen only on the loopback interface. There is no reason to expose the raw port if you are going to put a reverse proxy in front of it. The proxy will talk to Ollama locally, and the outside world will talk only to the proxy.
On Linux, Ollama runs as a systemd service. The OLLAMA_HOST variable is configured through a service override, which survives package updates.
- 01Edit the service overrideOpen the systemd override editor for the ollama service. This creates a clean drop-in file without touching the original service.
- 02Force local listeningSet OLLAMA_HOST to 127.0.0.1:11434. The daemon will then refuse any connection from the network.
- 03Reload and restartReload the systemd configuration, then restart the service to apply the variable.
- 04ConfirmDouble-check with ss -tlnp | grep 11434: the address should be 127.0.0.1, not 0.0.0.0.
#Step 2 — Add authentication through a reverse proxy
Since Ollama cannot authenticate, we delegate that responsibility to a reverse proxy placed in front of it. The proxy requires an identifier, checks the traffic, then relays it locally to Ollama. Two proven options: Caddy (minimal configuration, automatic TLS) and nginx (ubiquitous, extensively documented).
#Option A — Caddy (recommended for simplicity)
Caddy handles TLS automatically through Let's Encrypt and provides basic authentication in a few lines. First, generate a password hash, then reference it in the Caddyfile. Never store the password in plain text.
With a domain name pointing to your public IP and ports 80/443 open, Caddy obtains and renews the TLS certificate automatically. The client must provide the identifier with every request via the Authorization header.
#Option B — nginx
nginx requires a separate password file generated with htpasswd, plus a server block that applies authentication and proxies to Ollama. It's more verbose but very common, especially when nginx already serves other services.
#Step 3 — Encrypt traffic (TLS)
Basic authentication transmits the encoded identifier in base64: without TLS, it travels virtually in plaintext and is trivial to capture. Transport encryption is therefore not optional as soon as you leave localhost. There are two cases depending on whether you have a public domain name.
- Public domain + ports 80/443
- Let's Encrypt via Caddy (automatic) or certbot for nginx. Recognized certificate, no browser warnings, automatic renewal.
- Internal network without a public domain
- Self-signed certificate or internal certificate authority (mkcert). Clients will need to trust the certificate, but traffic remains encrypted on the LAN.
- Behind a VPN
- The VPN tunnel already encrypts everything. TLS is still recommended as defense in depth, but it’s less critical since no one external can reach the proxy.
#Step 4 — Clean remote access: VPN and Tailscale
The most important question: do you really need to expose Ollama to the internet? In the vast majority of cases, no. You want to access it from your own devices, not from the open web. A private network (VPN) meets that need exactly without ever exposing the port publicly.
Tailscale is the simplest option: it creates an encrypted mesh network (WireGuard) between your machines, with stable private IPs. Ollama then listens only on the Tailscale interface, and only your devices authenticated in your tailnet can reach it. No port forwarding and no public IP exposed.
- 01Install Tailscale on the serverInstall the client and connect the machine to your tailnet. It receives a private IP in 100.x.y.z, accessible only from your other authenticated devices.
- 02Link Ollama to the Tailscale interfaceSet OLLAMA_HOST to the machine's Tailscale IP (or keep 127.0.0.1 and expose it via Tailscale Serve). The port is reachable only within the tailnet.
- 03Install Tailscale on your clientsYour other machines join the same tailnet and reach Ollama through its 100.x.y.z IP address wherever you are, without opening a single port on the router.
- 04Verify isolationFrom an external network outside the tailnet, the port must be completely unreachable. This is the expected behavior.
If you insist on a traditional VPN, self-hosted WireGuard delivers the same result with more control: you create the tunnel, bind Ollama to the wg0 interface, and the port remains invisible from the internet. The choice between Tailscale and plain WireGuard mainly comes down to the tradeoff between simplicity and complete infrastructure sovereignty.
#Additional network hardening
Beyond the proxy and VPN, a few defense-in-depth measures limit the damage from misconfiguration. The guiding principle: never depend on a single layer.
- Strict firewall
- Block port 11434 inbound on all interfaces except loopback and VPN. With ufw: deny 11434 by default, allowing access only from the Tailscale/WireGuard subnet.
- No port forwarding
- Never create port forwarding for 11434 on your box/router. If you have one “for testing,” remove it: it is the n°1 cause of exposed instances.
- Rate limiting
- On the reverse proxy, apply rate limiting to mitigate potential abuse even after authentication (nginx limit_req, Caddy rate_limit).
- Strong passwords and rotation
- Basic authentication is only as strong as the secret. Use long passwords and change them if a client machine is compromised.
- Logging
- Enable the proxy's access logs to spot abnormal attempts. A healthy instance receives only your requests.
#Troubleshooting
- 403 Forbidden behind the proxy
- Ollama rejects the Host header. Force Host to localhost:11434 on the proxy side, or set OLLAMA_ORIGINS to allow your domain.
- Streaming freezes or arrives all at once
- The proxy buffers the response. Disable buffering (proxy_buffering off in nginx) and increase the read timeout for long generations.
- curl works, but not from another machine
- Ollama is still listening on 127.0.0.1 while the proxy is elsewhere, or the firewall is blocking the proxy’s port 443. Check ss -tlnp and the ufw rules.
- Let's Encrypt certificate fails
- Port 80 must be reachable from the internet for HTTP-01 validation, and the domain must point to the correct IP. Check the DNS and that port 80 is open.
- Tailscale: Ollama unreachable in the tailnet
- Ollama listens on 127.0.0.1 without tailscale serve. Either bind OLLAMA_HOST to the 100.x IP, or expose it through tailscale serve 11434.
- Still visible on Shodan after remediation
- The index takes time to refresh. First, confirm yourself from an external network that the port is closed; the index will update afterward.
#Go further
Securing access is only part of the picture. These guides build on this one with deployment and compliance:
- Deploy an AI chatbot for your team on the intranet
- The typical use case where this hardening applies: Ollama + Open WebUI multi-user deployments behind an authenticated reverse proxy.
- Deploy an LLM in production with Docker Compose
- Complete Ollama + Traefik stack, with the reverse proxy and network isolation managed from the outset.
- Local LLMs and GDPR: private-data compliance for businesses
- The regulatory counterpart: uncontrolled exposure is also a personal-data leakage risk that must be documented.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.