An offline AI: using an LLM without a connection internet
A local LLM only needs the internet to be downloaded. Once the model is on your disk, an offline AI works entirely without a connection: on a plane, in an area without coverage, while hiking, or on a workstation isolated from the network. This guide shows how to prepare your machine before leaving, which compact models to take with you, and how to verify everything so you aren't stuck at the worst possible moment.
#Why use offline AI
Cloud assistants (ChatGPT, Claude, Gemini) stop dead as soon as the connection drops. An offline AI, by contrast, runs entirely on your machine: the model is loaded into memory and computes your responses locally, without ever querying a remote server. It is the same principle as a conventional local LLM, taken to its logical conclusion—zero network dependency once setup is complete.
The situations where this changes everything are more common than you might think: a long-haul flight, a train trip without 4G, an assignment in an area without coverage, a construction site or industrial facility without Wi-Fi, or a workstation intentionally disconnected from the internet for privacy reasons. In all these cases, a well-prepared offline AI remains fully operational.
- Complete autonomy
- No interruption when the network disappears: on a plane, in a tunnel, in the countryside, or in a basement.
- Privacy
- Nothing leaves the machine. Ideal for sensitive data or an air-gapped workstation.
- Zero recurring cost
- No subscription, no per-token billing, no quota.
- Stable latency
- Speed depends only on your hardware, not on an overloaded server on the other side of the world.
#What needs the network (and what doesn't)
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Understanding the boundary between "online" and "offline" helps avoid surprises. Only one step requires the internet: downloading. Everything else is local.
- Network required (once)
- Download the application (Ollama, LM Studio), download the models (GGUF files), and retrieve updates.
- 100% offline
- Load an existing model, chat, generate text or code, summarize a local document, and use RAG on your files.
#Prerequisites before you begin
An offline AI fits on an ordinary laptop. What matters is the RAM (or VRAM if you have a GPU) and disk space to store the models.
- Free disk space
- Plan for 5 to 30 GB depending on the embedded models. A 7B model in Q4_K_M weighs about 4.5 GB, and a 14B about 9 GB.
- RAM / VRAM
- 8 GB for a 3B, 16 GB for a comfortable 7B, 32 GB to target a 14B. Q4 VRAM benchmarks: 3B≈2 GB · 7B≈5 GB · 14B≈9 GB.
- GPU (optional)
- A GPU provides a major speedup but isn’t required. A MacBook Apple Silicon (M-series) or a PC with 16 GB of RAM is sufficient on CPU alone.
- Battery
- Inference uses the CPU/GPU: on a laptop running unplugged, expect reduced battery life. Prefer compact models on the go.
#Download everything before losing network access
This is the decisive step and the only one that requires internet access. Do it comfortably at home or at the office over Wi-Fi, a few hours before departure—a model several GB in size doesn’t like unstable connections. Here’s how to do it with Ollama, the simplest approach for offline AI.
- 01Install Ollama while you have network accessDownload and install Ollama from ollama.com. This is the daemon that runs your models locally, listening on http://localhost:11434. Install it first, because the app itself is downloaded online.
- 02Download one or two compact modelsUse ollama pull to download the models to disk. Choose sizes suited to your machine (see the next section). A pull over stable Wi-Fi takes anywhere from a few minutes to a quarter of an hour, depending on the bandwidth and size.
- 03Install a graphical interface (optional)If you prefer a chat window to a terminal, install LM Studio (all-in-one, also handles downloads) or Open WebUI. LM Studio is particularly convenient on the go: app + models + interface in a single tool.
- 04Warm up each model onceStart a test conversation with each downloaded model while you still have network access. This confirms that the file is complete and loads correctly into memory.
#Compact models that are sufficient for mobility
When working offline on a laptop, efficiency comes first: a small, fast model that uses little battery is better than a behemoth that crawls and drains the battery. For most mobile use cases—writing, summarization, general questions, coding assistance—a 2B to 9B model in Q4_K_M does the job very well.
- Qwen 3.5 2B
- ~1.9 GB. The lightest option: ideal for an old laptop or for preserving battery life. Multimodal, 256k context, Apache 2.0 license. Sufficient for drafting, rewriting, and answering simple questions.
- Qwen 3.5 9B
- ~6.6 GB. Excellent all-rounder, strong at reasoning, vision, and coding, with 256k context. The best compromise for a 16 GB RAM laptop—the reference 8 GB choice in 2026.
- Granite 4.2 8B
- ~5.3 GB. Lightweight and token-efficient (IBM, Apache 2.0), 128k context, concise responses. A reliable choice for writing that preserves battery life.
- Granite 4.2 3B
- ~2.2 GB. Compact and nimble, highly efficient, designed for modest machines while remaining capable.
Q4_K_M quantization is the recommended default: it greatly reduces the model size and required memory, with minimal, imperceptible quality loss for most use cases. If you have RAM/VRAM to spare and want a little more precision, Q5_K_M is an alternative. Reserve Q8_0 and FP16 for very well-equipped machines.
#Verify that everything really works offline
Never take claims on faith: test in real-world conditions before you find yourself without a network connection. A three-minute test prevents disappointment at 10,000 feet.
- 01Cut off the network for goodTurn on airplane mode, disable Wi-Fi AND Ethernet. The goal is to exactly reproduce a complete lack of connectivity.
- 02Start a conversationOpen your terminal or LM Studio and chat with each downloaded model. If the model responds, your offline AI is operational.
- 03Check that no network error is blocking accessSome interfaces display an offline warning banner: as long as generation works, ignore it. If the chat refuses to start, the model probably wasn’t fully downloaded.
#Isolated workstation and air-gapped use
Some machines are deliberately kept permanently offline: workstations processing confidential data, industrial environments, and laboratories. This is the “air-gapped” use case—the target machine will never have network access, including for installation.
The principle: prepare the files on a connected machine, then transfer them via USB drive or external disk to the isolated workstation. Ollama stores its models in a directory that can be copied as is.
- Offline installers
- Retrieve the installation binary for Ollama or LM Studio from a connected machine, then copy it to the isolated workstation.
- Model files
- On Linux/macOS, Ollama models live in ~/.ollama/models; on Windows, in %USERPROFILE%\.ollama\models. Copy this folder to the target machine in the same location.
- GGUF Alternative
- With LM Studio, simply copy the downloaded .gguf files into the app’s model folder on the isolated machine.
#Update your models when the network comes back
Once you're back in a covered area, it's time to maintain your setup: download new model versions, add new models you discovered along the way, and update the application. None of this can be done offline, which is why it's important to do it as soon as the network returns.
- 01Refresh existing modelsA ollama pull sur model that is already present downloads only the modified layers when a new version is available. It's fast and keeps your models up to date.
- 02Update the applicationRerun the Ollama or LM Studio installer to get bug fixes and support for new models. Updates are frequent in this ecosystem.
- 03Clean upDelete models you no longer use to free disk space before your next trip.
- 04Re-test in airplane modeAfter every update, run an offline test: a new app version could introduce unexpected behavior without network access.
#Troubleshooting
- “Model not found” offline
- The model was not (completely) downloaded before the interruption. Reconnect, rerun ollama pull jusqu'until the success message, then verify with ollama list.
- The app refuses to start without a network connection
- Some interfaces attempt an online check at startup. Wait a few seconds for it to time out, or disable update checks in the settings.
- Very slow responses while traveling
- On pure CPU and battery power, this is normal. Switch to a smaller model (3B) and close resource-intensive applications. Plug in the laptop if possible.
- Battery draining before your eyes
- Inference puts a heavy load on the processor. Reduce the model size, limit response length, and enable power-saving mode for non-urgent tasks.
- Copied model that won't load (air gap)
- The copy was incomplete: blobs or manifests were missing. Copy the entire .ollama/models folder again.
- Not enough memory
- The model exceeds your RAM/VRAM. Choose a smaller size or a lighter quantization (Q4_K_M rather than Q8_0).
#Go further
Offline AI relies on the same foundations as a conventional local installation. These site guides complete the preparation:
- Install Ollama
- The step to perform while you still have network access to set up the inference engine.
- Run an LLM locally without a GPU (CPU only)
- Essential for mobility: expected models and speeds when running on battery in pure CPU mode.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To balance disk size, memory, and quality based on the machine you carry with you.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.