Share Ollama on your local network (family, team)
Running Ollama as a local network server changes everything: one GPU, multiple users. Your partner on their MacBook, the developer at their desktop, the child on the tablet—everyone uses the same daemon without duplicating 30 GB of models. This guide shows how to expose Ollama properly on the local network without opening it to the entire world.
#Why run a Ollama server on a local network?
By default, Ollama listens only on 127.0.0.1:11434—so only the machine it runs on can access it. That’s secure, but it wastes a GPU. A RTX 4090 or a Mac Studio M4 Max can handle 3 to 5 simultaneous users comfortably with 7B–14B Q4 models.
- Pool VRAM
- A single model loaded in memory for the whole household or team—no redundant copies.
- Centralize models
- 150 GB of GGUF stored once on the server, never on the laptops.
- Standardize versions
- Everyone runs the same Qwen 3.5 9B Q4—no more quality differences between workstations.
- Save battery
- The laptops do no computing—they send an HTTP POST to the fixed server.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Ollama installed
- On the machine that will host the server (Linux, macOS, or Windows). Ideally, the machine with the best GPU.
- Static IP or local DNS
- The server must keep the same IP. DHCP reservation on the router or a static IP. Otherwise, hostname.local via mDNS.
- Trust network
- Home Wi-Fi with WPA2/3, or a team VLAN. No shared public Wi-Fi.
- Admin rights
- To modify the firewall and the system service (systemd, launchd, services.msc).
#1. Expose Ollama with OLLAMA_HOST=0.0.0.0
Ollama reads two key environment variables: OLLAMA_HOST defines the listening interface, and OLLAMA_ORIGINS defines the allowed CORS origins. To switch to server mode, change OLLAMA_HOST from 127.0.0.1 to 0.0.0.0 (all interfaces).
#Linux (systemd)
On modern distributions, Ollama runs through systemd. Edit the service override rather than the base file — it survives package updates.
In the editor that opens, paste this block between the commented lines:
Check that ss -tlnp actually shows Ollama on 0.0.0.0:11434, rather than only on 127.0.0.1:11434.
#Windows
On Windows, Ollama reads user environment variables at startup. The proper method: add OLLAMA_HOST to the user variables, then restart the service from the taskbar (right-click the icon → Quit Ollama, then restart).
#2. Open the firewall for the LAN
Once Ollama is listening on 0.0.0.0, the firewall must allow port 11434—but only from the local network. Never open 11434 to the Internet: Ollama has no native authentication.
#UFW (Ubuntu, Debian)
Adapt 192.168.1.0/24 to your actual subnet (use ip a to check). If you leave ufw allow 11434 unrestricted by source and the machine is exposed behind port forwarding, you are offering Ollama to the entire world—anonymously.
#iptables (without UFW)
#Windows Defender Firewall
#3. macOS specifics
On macOS, Ollama runs as a GUI app (an icon in the menu bar) that launches the daemon in the background. Environment variables defined in the shell are not visible to the GUI app—you need to use launchctl or modify the app.
- 01Exit Ollama completelyClick the llama icon in the menu bar → Quit Ollama. Run ps aux | grep ollama to verify that nothing is still running.
- 02Define the variable at the launchctl levelIn a terminal: launchctl setenv OLLAMA_HOST "0.0.0.0:11434" then launchctl setenv OLLAMA_ORIGINS "*". These variables are inherited by all GUI apps launched afterward.
- 03Restart Ollama.appOpen Applications → Ollama.app. The app will reread the environment and listen on 0.0.0.0.
- 04Persist across rebootslaunchctl setenv doesn’t survive a reboot. To make it permanent, create a LaunchAgent at ~/Library/LaunchAgents/com.ollama.env.plist (see the Apple documentation) or rerun the setenv commands from a startup script.
#4. Connect the clients
From another machine on the LAN, replace localhost with the server's IP in all commands and configurations.
For the ollama CLI itself, export OLLAMA_HOST on the client side—the ollama run pointe command then runs against the remote server.
- Open WebUI
- In Settings → Connections, add the http://192.168.1.42:11434 URL as the Ollama API URL. The interface runs on any machine on the LAN.
- Continue.dev (VS Code)
- In config.json, the apiBase field of the ollama provider points to http://192.168.1.42:11434.
- LangChain Python
- Ollama(base_url="http://192.168.1.42:11434", model="qwen3.5:9b") — identical to localhost locally.
#5. Nginx reverse proxy + basic authentication
Exposing Ollama raw on the LAN is fine for the family, but for a team it’s better to add at least HTTP Basic authentication and some logging. Nginx does this in 20 lines.
Important: with this setup, Ollama must point back to 127.0.0.1:11434 (not 0.0.0.0). Nginx listens publicly on port 8080, checks the IP and password, then forwards to Ollama locally. Clients now connect to http://alice:motdepasse@192.168.1.42:8080.
#Security best practices
- Restrict by source IP at the firewall
- Belt and suspenders: even with Nginx, keep the UFW/iptables rule. An Nginx bug or an incorrect bind can expose 11434 directly.
- Disable /api/create externally
- This route allows any arbitrary Modelfile to be pushed. With Nginx, add location /api/create { return 403; }.
- Log requests
- Nginx access_log shows who is calling what. Useful for spotting a server IP leak or a misconfigured client that is spamming.
- Per-user quotas
- Not native to Ollama. To limit Bob from launching a large generation job at 3 a.m., consider LiteLLM or Open WebUI as a middleware layer.
- Back up ~/.ollama
- The centralized server becomes a single point of failure. The models directory can weigh 100+ GB—at a minimum, use an rsync script to an external drive.
#Troubleshooting
- Connection refused from a client
- Ollama still listens on 127.0.0.1. Check ss -tlnp | grep 11434 on the server—if it shows 127.0.0.1:11434, the OLLAMA_HOST variable isn’t visible to the service. Recheck systemctl show ollama | grep Environment.
- Connection timed out
- The firewall is blocking it. Run nc -zv 192.168.1.42 11434 from the client: if it times out, it is the firewall. If it is refused, Ollama is not listening.
- CORS error in Open WebUI
- OLLAMA_ORIGINS=* is missing or was not applied. For debugging: curl -H "Origin: http://autre-machine" -I http://192.168.1.42:11434 — the response should contain Access-Control-Allow-Origin.
- macOS: variable ignored after restart
- launchctl setenv doesn’t persist. You need a LaunchAgent. Pragmatic alternative: a ~/start-ollama.sh script that runs launchctl setenv + open Ollama.app, launched manually after each reboot.
- Sudden slowdown with multiple users
- Ollama processes queued requests by model. With 3 simultaneous users on a 14B, the 3rd user waits. OLLAMA_NUM_PARALLEL=2 (env var) allows 2 requests in parallel at the cost of more VRAM.
- The model is unloaded between requests
- By default, Ollama unloads a model 5 minutes after the last request. OLLAMA_KEEP_ALIVE=24h forces it to stay in VRAM—crucial for a shared server.
#Go further
The shared Ollama server is the foundation of a multi-user setup. Three natural next directions:
- Add a shared chat interface
- The Open WebUI guide with Ollama: the complete guide details setting up a multi-user ChatGPT-like frontend that connects to this server.
- Deploy to production with Docker
- The Deploying an LLM in Production with Docker Compose guide shows the same containerized architecture with Traefik and HTTPS.
- Extend it to an intranet chatbot
- The guide Deploying an AI chatbot for your team on an intranet adds SSO authentication, monitoring, and conversation backups.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.