Air-gap your machine of inference
Air-gapping an inference machine means cutting it off from all network connections (wired, Wi-Fi, and Bluetooth) and allowing models and updates to enter only through controlled physical media. An outbound firewall is a useful layer, but it isn’t an air gap: it’s only a software configuration. The real weak points are software that calls out (Ollama Cloud, updates, downloads) and the USB drive.
The word “local” is not enough for some sensitive contexts: a local machine remains connected, and its software may contact the outside world. This guide distinguishes true air-gapping from firewall-based isolation, lists the components of an LLM stack that call out by default and how to disable them, describes transferring models by USB drive with fingerprint verification, and proposes leak-proofing tests you can rerun.
#When an air gap is justified—and when it costs too much
An air gap should be decided through a risk analysis, not as a matter of principle. Cases where it is justified include classified data or data covered by a high-level trade secret, contractual or sector-specific requirements that explicitly mandate an isolated network, environments that already use isolation and where the LLM must be integrated, and workstations dedicated to exceptionally sensitive files (a lawyer, a researcher). For ordinary personal data, including health or financial data, regulations call for measures proportionate to the risk rather than total isolation: a local LLM on an encrypted workstation, without a cloud account, already meets most needs. The guide to local AI in the enterprise and the privacy checklist cover this middle ground.
Before you get started, ask three questions. Who will handle the transfers, and how often? Which components does the stack require us to update (Ollama, llama.cpp, the interface, the models)? And what happens when a newer model becomes essential: the process must remain practical without exceptions, because an exception made once becomes the rule. If you cannot answer these three questions, start with an offline workstation per use case and a strict firewall, and move up a level only when the risk justifies it.
#Air gap, near-air gap, and firewall: don't confuse them
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
| Tier | What it is | What remains possible |
|---|---|---|
| Machine kept offline by design | No active connection, but the network card is present | A mistaken or software-triggered reconnection; it all depends on discipline |
| Strict firewall (deny by default, including outbound traffic) | Software configuration that blocks outbound traffic | A rule added “just to try it,” a system flaw, a service that bypasses |
| Physical air gap | No active network interface: cable disconnected, Wi-Fi and Bluetooth disabled in firmware or removed | Leaks through removable media; physical channels (outside this guide's scope) |
This guide covers all three: the firewall is the minimum layer every sensitive workstation should have, and a physical air gap is the top layer. On a Mac, macOS-specific settings belong in the guide to optimizing Macs with Apple chips; under Windows, the principles are the same, but control over components that call out is less granular than with a minimal Linux system, hence the preference here for Linux.
#Prepare the machine
- Minimum system
- Debian or Ubuntu Server without a graphical interface for an inference server. No browser, no package manager connected to an app store.
- Automatic update services
- Disable update timers, for example sudo systemctl disable --now apt-daily.timer apt-daily-upgrade.timer, so that no connection attempt is made silently.
- Accounts
- An account dedicated to LLM use, with sudo restricted to the administrator.
- Radios and interfaces
- Wi-Fi and Bluetooth disabled in the BIOS or firmware; better yet, physically absent. No network cable connected.
- No double startup
- A connected partition alongside it creates indirect paths (swap file, shared files).
- Time
- Set the clock manually or from an internal source: synchronization with a public time server is outbound traffic.
- Encrypted disk
- An isolated workstation can be stolen or seized: encrypt the disk (see the dedicated guide).
#LLM stack software that makes outbound calls by default
For a physically isolated machine, these calls fail harmlessly; for a machine protected only by a firewall, or one awaiting disconnection, they are the real risk. A typical local stack contains several components that contact the Internet unless prevented from doing so. Here are the ones the official documentation allows you to disable.
| Component | What it can do | Documented setting |
|---|---|---|
| Ollama | Hosted models and web search (Ollama Cloud) | OLLAMA_NO_CLOUD=1 or disable_ollama_cloud in ~/.ollama/server.json; the logs then show “Ollama cloud disabled: true” |
| Ollama (network exposure) | Local HTTP server | Listens on 127.0.0.1 by default; change OLLAMA_HOST only for an isolated internal network |
| Open WebUI | Update checks, downloading embedding models from Hugging Face | OFFLINE_MODE=true (also disables ENABLE_VERSION_UPDATE_CHECK); download the embedding models first |
| Hugging Face libraries (transformers, sentence-transformers) | Hub calls to verify or download | HF_HUB_OFFLINE=1: no HTTP calls to the Hub; only cached files are used |
| Vector databases | Potential telemetry | Weaviate: DISABLE_TELEMETRY=true; check the equivalent option for each component |
The Ollama documentation specifies that when using Ollama locally, the vendor sees neither the requests nor the data; disabling cloud features removes access to hosted models and web search, which is exactly the point. Be careful with Open WebUI: the documentation warns that if you have not downloaded an embedding model before enabling OFFLINE_MODE, RAG, web search, and document analysis features may not work. Download everything the machine will need before taking it offline.
#Strict firewall: deny by default, inbound and outbound
The firewall is the second line of defense, after physical disconnection, and the only protection on a workstation that remains on an internal network. The principle is deny by default in both directions, with a few explicit permissions to the isolated internal network. Write the allow rules before enabling the firewall; otherwise, you'll cut yourself off from an SSH session.
Traffic on the local interface (loopback) is required for Ollama and Open WebUI, which communicate over 127.0.0.1; check with ufw status verbose that your configuration does not block it. If the machine is on an internal LAN, add protection at the router as well: a VLAN with no route to the Internet.
#Check what Ollama is listening on
#Transfer models and updates by USB drive
An isolated machine receives models (from a few GB to several dozen), updates (Ollama, llama.cpp, Open WebUI), and sometimes documents. The critical point is the USB drive: it’s the only door, and therefore the only entry point for malware. Always use the same dedicated drive, format it regularly, and never use a drive that has been used elsewhere.
- 01Download on a connected machineRetrieve the model's GGUF file from Hugging Face or the update binary from the official website. Record the SHA-256 fingerprint displayed by the source when available.
- 02Calculate the fingerprint and analyzeRun sha256sum on the file, then pass it to an antivirus (such as ClamAV) on this transit machine, which is not the isolated machine.
- 03Copy to the USB driveCopy the file and a text file containing its hash.
- 04Check on the isolated side before useOn the isolated machine, recalculate the hash: it must be identical. A difference means a file was modified or corrupted: do not use it.
- 05Import into OllamaCreate a Modelfile with FROM /chemin/vers/modele.gguf and then ollama create nom-du-modele -f Modelfile. Ollama does not quantize a GGUF during import: the file must already have the desired quantization.
The alternative method is to copy the entire model directory from a connected machine (blobs and manifests) to the isolated machine. It works, but the directory is in a different location depending on the installation: the Ollama FAQ gives ~/.ollama/models on macOS, /usr/share/ollama/.ollama/models on Linux with the standard installer, and C:\Users\%username%\.ollama\models on Windows. Using an isolated GGUF file is easier to verify because there is only one fingerprint to compare. The guide on import GGUF provides details on the Modelfile.
#Harden the USB port
A USB key can masquerade as a keyboard or another device. On Linux, USBGuard can be described in one sentence as a USB device allowlisting tool: it defines which types of devices are authorized and how they can interact with the system. Also disable automatic media mounting so that nothing runs when connected. In highly sensitive contexts, a data diode (a device that allows data to pass in only one direction) replaces the USB key, at the cost of a more complex setup.
#Test isolation
A test proves airtightness only for the period it observed: repeat it after every update and let the capture run through a complete work session. The minimum protocol covers operation, no routing, no name resolution, and actual traffic.
- Offline operation
- Unplug the cable and turn off all radios, then run the full pipeline (model, RAG, interface). Everything should work; otherwise, a component depended on the Internet.
- No route
- ping -c 1 8.8.8.8 should fail (“Network is unreachable” or equivalent).
- No DNS
- host example.org must fail; check /etc/resolv.conf: no external server.
- Established connections
- sudo ss -tunp state established : aucune connexion vers une adresse hors du réseau interne.
- Real-world traffic
- A tcpdump capture on all interfaces, excluding the internal network, must remain empty throughout an entire usage session.
#Regular audit
| Frequency | Control | Objective |
|---|---|---|
| Monthly | ufw status verbose and list of rules; ss -tlnp | Identify a rule added “just to try it” |
| With every transfer | Fingerprints before and after the key; antivirus scanning | Detect a modified file |
| With every component update | Rerun the leak-tightness tests and reread the offline variables | A new component can add an outbound call |
| Semiannual or annual | Clean reinstallation from verified images; threat review | Eliminate an accumulated compromise |
#Frequently asked questions about air-gapping an LLM machine
Is an outbound firewall enough to qualify as an air gap?+
How do you update Ollama on a machine without Internet access?+
How do you transfer a model to the isolated machine?+
Does Ollama send my requests outside?+
Is the USB drive the weak link?+
Should you encrypt the disk of an already isolated machine?+
- Encrypt the model drive
- Privacy checklist
- Local AI in the enterprise: GDPR, sovereignty, and deployment
- Import a GGUF model from Hugging Face into Ollama
- Get the most out of a Mac Apple Silicon
- Source: Ollama FAQ, disableable cloud and locations
- Source: Ollama, import d's GGUF model
- Source: Open WebUI, environment variables
- Source: USBGuard
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.