Beginner 14 minPrivacy

Checklist privacy

Direct response

For a local LLM to remain private, check five points: the tool’s cloud features, library telemetry, actual network traffic, the origin of models and extensions, and synchronized or backed-up folders. The decisive test takes one action: disconnect from the Internet and run a normal session again. If everything works, the chain is local; otherwise, there is a call to identify.

A local LLM is confidential by design, but only if nothing around the model communicates externally: cloud option, library ping, plugin, synchronized folder. These leaks are rarely intentional and almost always invisible. This checklist gives you verifiable criteria, commands, and dated settings to audit your installation in fifteen minutes.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What “local” guarantees, and what it does not

To verify that a local LLM remains private, you need to check five points: the tool you use, its telemetry and calls to remote services, the machine's actual traffic, the provenance of models and extensions, and finally the folders where your conversations live. A model running on your machine does not send anything on its own: leaks almost always come from what surrounds it—an enabled cloud option, a library ping, a plugin, or synchronization. The most reliable method is still a network test: disconnect from the Internet, run a normal session again, then observe what still tries to get out. The rest of this page details each check, including the environment variables, commands, and settings to verify.

i
The firewall test
Disconnect the machine’s outgoing connection (airplane mode, unplugged cable, or firewall rule), then run a normal session again: chat, document import, and searching your notes. If everything works, your pipeline is local. If a tool complains or remains stuck, note which one: it may be a potential leak to identify, or a cloud feature you were using without realizing it.

#What level of control do you have over your data?

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Not everyone needs the same level of rigor. Personal use without sensitive data can settle for the first few points; a client or medical case requires complete isolation, and an industrial secret or NDA-covered setting requires proof of what leaves the system. The following table links the required level to the controls to check.

Recommended control level based on data sensitivity
Data typeMinimum checksAdd if an audit is possible
Typical personal use (ideas, writing, personal code)Verified tool, cloud features disabled, telemetry turned offTest airplane mode once
Customer, HR, healthcare, and legal dataEverything above, plus an application firewall, folders excluded from synchronization, and an encrypted driveFirewall log retained, list of allowed extensions
Trade secrets, code under an NDA, regulated dataEverything above, plus a machine isolated from the network for inferenceComplete air gap, models imported from a separate workstation, extension review

The last two lines are detailed in the guides on isolating an inference machine and encrypting the disk.

#1. Check the tool itself

Each inference application has its own policy. The following reflects what the vendors state in their documentation or privacy policy, reviewed on September 30, 2026; an update may change these settings, so check again after every major version upgrade.

What vendors say about data leaving the system
ToolWhat is indicatedCheck on your system
OllamaThe FAQ states that prompts and data are not viewed when the model runs locally; the service listens on 127.0.0.1, port 11434That you do not use a cloud model or web search
LM StudioThe privacy policy states that messages, histories, and documents do not leave your machine when using local models; only model searches or downloads and update checks contact the publisherPaid cloud features (hosted models, web search) remain disabled
JanAdvertised as 100% offline, with data stored locally; remote APIs (OpenAI, Groq) send your data to their serversThat no remote API key is configured
GPT4AllA Datalake setting, which you must enable intentionally, lets you share interactions with the community; it is disabled by defaultThat Datalake is indeed turned Off
Open WebUIThe repository provides variables for disabling library analytics and telemetry (see the next section)No remote provider key as a fallback

Two corrections to common misconceptions. First, the “opt-in telemetry” wording often seen for LM Studio is outdated: its current policy states that the application does not include telemetry. Second, saying that Ollama has “no remote functionality” has been inaccurate since cloud models and web search were introduced. These features are optional, but present.

#Ollama: disable cloud features

If Ollama is your engine, explicitly disable anything that could leak data. The FAQ states that disabling the cloud means losing the ability to use cloud models and Ollama's web search. That's exactly the desired effect for a confidential workstation.

Disable the Ollama cloud (Linux, systemd service)
sudo systemctl edit ollama
# Ajouter dans l'éditeur :
# [Service]
# Environment="OLLAMA_NO_CLOUD=1"
sudo systemctl restart ollama

On macOS with the app, set the variables with launchctl setenv before relaunching the app, as described in the FAQ; alternatively, add the disable_ollama_cloud key set to true in ~/.ollama/server.json. After restarting, the logs show "Ollama cloud disabled: true", providing simple evidence to keep.

#2. Library telemetry and analytics

The tool itself may be clean while the libraries it bundles talk. In the Hugging Face Python ecosystem, the documentation states that some data is collected by default to monitor usage and prioritize features, with the central implementation in huggingface_hub. One variable disables them all.

Disable Hugging Face telemetry
export HF_HUB_DISABLE_TELEMETRY=1
# Équivalent reconnu par tout l'écosystème (transformers, diffusers, gradio) :
export DO_NOT_TRACK=1

# À placer dans ~/.zshrc ou ~/.bashrc pour la rendre persistante

The HF_HUB_DISABLE_IMPLICIT_TOKEN variable, often mentioned alongside the previous ones, is not a telemetry variable: it prevents your Hugging Face token from being sent automatically with requests that do not need it, limiting what you expose while signed in. It is useful, but for a different reason, and the documentation warns that it can have side effects (private models no longer appear in some lists).

Force offline mode once the model is cached
export HF_HUB_OFFLINE=1
# Aucun appel HTTP vers le Hub ; seuls les fichiers du cache sont lus

This last setting is the most powerful option for a sensitive workstation: it turns a silent leak into a visible error. If a script fails and reports that the cache is missing, it was trying to reach the Hub—which is what you wanted to know. The old TRANSFORMERS_OFFLINE variable served the same purpose for a long time, but HF_HUB_OFFLINE is the one highlighted by the current Transformers documentation.

#Open WebUI and its dependencies

Open WebUI includes third-party components, including a vector database and an installation-statistics module. Its sample configuration file groups three lines in a “DO NOT TRACK” section that disable these transmissions. They can be set through environment variables, in the configuration file, or in docker run.

Anti-telemetry variables for Open WebUI
SCARF_NO_ANALYTICS=true
DO_NOT_TRACK=true
ANONYMIZED_TELEMETRY=false

Also check that you have not added a remote provider key: a “fallback” to a hosted service, when the local model is unavailable, silently sends your prompts outside. The Open WebUI guide with Ollama shows where these connections are configured.

#3. Network monitoring: see what is really leaving the system

Vendor claims can be read; traffic can be measured. An application firewall lists every program that attempts an outbound connection and lets you decide case by case, revealing within days what no documentation says.

Little Snitch (macOS)
Per-application firewall: see who is calling which domain and decide on a case-by-case basis.
GlassWire (Windows)
Visual history of traffic by process, with alerts for new connections.
OpenSnitch (Linux)
Interactive application firewall for GNU/Linux, inspired by Little Snitch, open source (GPL-3.0 license).
ufw or iptables
They block by default and allow only what is necessary; the air-gap guide details the process.

#List connections and listening ports

Two separate commands answer two separate questions. The first shows which address the service accepts connections on; the second shows open connections to the outside.

Listening and connections from Ollama (Linux and macOS)
# Qui écoute, et sur quelle adresse ?
lsof -nP -iTCP -sTCP:LISTEN | grep -i ollama

# Quelles connexions sortantes sont ouvertes ?
lsof -nP -i | grep -i ollama | grep ESTABLISHED

Read the result as follows: listening on 127.0.0.1:11434 is the default behavior, documented in the Ollama FAQ, and it is reachable only from the machine itself. Listening on * or 0.0.0.0 means the service is reachable from the local network, which happens when the OLLAMA_HOST variable has been changed. This is not an Internet leak, but it is an exposure that must be addressed: see the guide on securing a Ollama server. The commonly misunderstood point: 0.0.0.0 does not mean “localhost only.”

An ESTABLISHED connection to a public address during inference is the signal to investigate. Note that downloading a model legitimately generates connections; it’s inference, once the model is in place, that should remain silent. Application updates may also contact the vendor, as LM Studio acknowledges in its policy: decide whether you’re willing to allow them.

#4. Models: where they come from and what they contain

The first model download requires a connection: that is normal, and what is transferred then is the incoming model, not your data. Once the file is in place, a model is a set of weights: it does not open a connection by itself. The key concern is the format and what runs it.

Weight format
GGUF and safetensors formats contain only tensors. The pickle format, used by default for PyTorch weights, is different: the Hugging Face documentation warns that loading a pickle file can execute arbitrary code.
Origin
Prefer publishers’ repositories (Mistral AI, Meta, Qwen, Google) or the libraries of a tool that verifies its models. An anonymous repository in pickle format is one to avoid.
What runs
The risk comes from the application and its tools: if the model can call a web search, an MCP server, or a script, those tools access the network, not the weights.
!
A model with tools can output data
A model alone can exfiltrate data only through its output text, which you read. As soon as you connect tools to it—browsing, MCP, code execution—it can ask the application to send a request containing your data. Treat every tool as an extension, with the controls described in the next section.

#5. Extensions, plugins, and connected tools

This is where most real leaks occur. A code editor extension, note-taking plugin, tool declared in Open WebUI, or MCP server runs with your permissions and can call any domain. A plugin that promises “local AI” is only as trustworthy as its author.

Read before enabling
Search the code for network calls (HTTP requests, tracking libraries) and the endpoint actually configured: it is often a single line to change to point to Ollama instead of a hosted service.
Prefer open source
Code you can read is verifiable; an assistant with a proprietary backend is not.
Limit permissions
A tool that does not need to read the entire disk should not have access to it; give it a dedicated working directory.
Monitor with the firewall
After installation, let the application firewall learn which domains the extension contacts for a week.

#6. Storage, backups, and synchronization

Your conversations and indexed documents are files. What happens to them depends on where the application stores them and what monitors that folder: cloud synchronization, online backup, or a search indexer.

Where are the Ollama models?
According to the FAQ: ~/.ollama/models on macOS, /usr/share/ollama/.ollama/models on Linux, C:\Users\%username%\.ollama\models on Windows, movable with OLLAMA_MODELS. These files are several gigabytes in size, so excluding them from synchronization also saves bandwidth.
Where conversations are stored
In the data directory of the interface you use (LM Studio, Jan, GPT4All, Open WebUI), not the one for Ollama. Locate this directory in the application settings, then verify that no cloud service is monitoring it.
Synchronization
Many applications store their data in Documents, which iCloud, OneDrive, or Dropbox automatically replicate. A confidential conversation there becomes a copy on a third-party server.
Backups
An off-site backup or an unencrypted external drive copies your history. Encrypt the destination, or exclude the folder.
Disk
An unencrypted drive makes every conversation readable if stolen: FileVault, BitLocker, or LUKS, with a strong passphrase.

#15-minute audit: the procedure

This procedure runs the checks in the order that reveals the most with the least effort. Keep a screenshot or one log line for each step: it serves as evidence when dealing with a client.

  1. 01
    Inventory the stack
    List every program that touches your prompts: engine (Ollama, LM Studio), interface (Open WebUI, Jan), editor extensions, and Python scripts. Anything not listed is not controlled.
  2. 02
    Disable remote features
    Disable cloud models and web search (OLLAMA_NO_CLOUD=1), remove remote API keys, and verify that Datalake is set to Off in GPT4All.
  3. 03
    Set the variables
    Put HF_HUB_DISABLE_TELEMETRY=1, DO_NOT_TRACK=1, and, once the models are cached, HF_HUB_OFFLINE=1 in your shell profile; add the three anti-analytics lines to Open WebUI.
  4. 04
    View network
    Install an application firewall, run a normal ten-minute session, then review the domains contacted. Also run the two lsof commands.
  5. 05
    Move folders
    Move conversations and models out of synchronized folders, and encrypt the disk and the backup destination.
  6. 06
    Replay in airplane mode
    Disconnect and run the same session again. Anything that fails is either a leak or a cloud function to remove; fix it, then start over until the test passes.

#The summary checklist

☐ Cloud features disabled
OLLAMA_NO_CLOUD=1, no hosted models, no remote API key, Datalake Off.
☐ Telemetry disabled
HF_HUB_DISABLE_TELEMETRY=1 and DO_NOT_TRACK=1; three anti-analytics variables in Open WebUI.
☐ Application firewall active
Little Snitch, GlassWire, OpenSnitch, or ufw with a strict rule and reviewed logs.
☐ Listening verified
127.0.0.1 for Ollama, unless deliberate and secured exposure is intended.
☐ Offline models
After downloading, HF_HUB_OFFLINE=1; weights in GGUF or safetensors, not an unknown pickle.
☐ Extensions reviewed
Code read, endpoints verified, permissions limited.
☐ Folders not syncing
Conversations and models excluded from iCloud, OneDrive, Dropbox, and unencrypted backups.
☐ Encrypted drive
FileVault, BitLocker, or LUKS.
☐ Airplane mode test passed
A complete session without the Internet.

#Frequently asked questions about the privacy of a local LLM

FAQ
Does a local LLM send my data to the vendor?+
Not the model itself: it only performs calculations. Ollama states that it does not see your prompts during local use, and LM Studio specifies that messages, histories, and documents do not leave your machine when using local models. Ancillary features remain: cloud models, web search, updates, and plugins. Check them one by one.
How can you verify that Ollama does not connect to anything?+
Run lsof -nP -iTCP -sTCP:LISTEN to see the listening address (127.0.0.1:11434 by default), then lsof -nP -i during inference to identify any ESTABLISHED connections. Add an application firewall to keep a history of connection attempts, and use OLLAMA_NO_CLOUD=1 to disable remote features. If a connection appears, identify its domain before drawing conclusions.
How do you disable Hugging Face telemetry?+
Set HF_HUB_DISABLE_TELEMETRY=1 in your environment variables. DO_NOT_TRACK=1 has the same effect across the entire ecosystem (transformers, diffusers, gradio). To block all calls to the Hub once the models are cached, add HF_HUB_OFFLINE=1. These variables are read when the library is imported: set them before launching Python.
Do you need an additional firewall if the tool says it sends nothing?+
For personal use, it is not essential. For customer data or data under an NDA, yes: a vendor’s statement is not proof, and an update, plugin, or library may contact a server. An application firewall turns trust into an observable fact, and its log provides a record.
Are my conversations synced with iCloud or Dropbox?+
It depends on where the application stores its data folder. If it is under Documents or in a synchronized folder, yes. Find the folder in the application settings, then exclude it from synchronization or move it to a non-replicated location. Do the same with backups, which copy the history, and encrypt their destination.
Is airplane mode enough to prove that my setup is local?+
This is the best simple test: if your normal session works without a network connection, no essential component depends on a remote service. It does not prove that there was no attempt to send data silently, which would fail without raising an error. Supplement it with an application firewall that logs attempts, and review the log after a week of real-world use rather than after a single test.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.