Beginner 8 minOffline

An offline AI: using an LLM without a connection internet

A local LLM only needs the internet to be downloaded. Once the model is on your disk, an offline AI works entirely without a connection: on a plane, in an area without coverage, while hiking, or on a workstation isolated from the network. This guide shows how to prepare your machine before leaving, which compact models to take with you, and how to verify everything so you aren't stuck at the worst possible moment.

By Léa B.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why use offline AI

Cloud assistants (ChatGPT, Claude, Gemini) stop dead as soon as the connection drops. An offline AI, by contrast, runs entirely on your machine: the model is loaded into memory and computes your responses locally, without ever querying a remote server. It is the same principle as a conventional local LLM, taken to its logical conclusion—zero network dependency once setup is complete.

The situations where this changes everything are more common than you might think: a long-haul flight, a train trip without 4G, an assignment in an area without coverage, a construction site or industrial facility without Wi-Fi, or a workstation intentionally disconnected from the internet for privacy reasons. In all these cases, a well-prepared offline AI remains fully operational.

Complete autonomy
No interruption when the network disappears: on a plane, in a tunnel, in the countryside, or in a basement.
Privacy
Nothing leaves the machine. Ideal for sensitive data or an air-gapped workstation.
Zero recurring cost
No subscription, no per-token billing, no quota.
Stable latency
Speed depends only on your hardware, not on an overloaded server on the other side of the world.

#What needs the network (and what doesn't)

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Understanding the boundary between "online" and "offline" helps avoid surprises. Only one step requires the internet: downloading. Everything else is local.

Network required (once)
Download the application (Ollama, LM Studio), download the models (GGUF files), and retrieve updates.
100% offline
Load an existing model, chat, generate text or code, summarize a local document, and use RAG on your files.
i
The classic trap
Many tools attempt an update check or an automatic “pull” at startup. Without a network connection, this isn’t blocking as long as the model is already downloaded—but the app may display a misleading error. Always test in airplane mode BEFORE you leave.

#Prerequisites before you begin

An offline AI fits on an ordinary laptop. What matters is the RAM (or VRAM if you have a GPU) and disk space to store the models.

Free disk space
Plan for 5 to 30 GB depending on the embedded models. A 7B model in Q4_K_M weighs about 4.5 GB, and a 14B about 9 GB.
RAM / VRAM
8 GB for a 3B, 16 GB for a comfortable 7B, 32 GB to target a 14B. Q4 VRAM benchmarks: 3B≈2 GB · 7B≈5 GB · 14B≈9 GB.
GPU (optional)
A GPU provides a major speedup but isn’t required. A MacBook Apple Silicon (M-series) or a PC with 16 GB of RAM is sufficient on CPU alone.
Battery
Inference uses the CPU/GPU: on a laptop running unplugged, expect reduced battery life. Prefer compact models on the go.
→
It still works without a GPU
On a plane or train, you will often be running on CPU only. A 3B or 7B model in Q4_K_M remains responsive on a recent laptop. The “LLM locally without a GPU” guide details the models and expected speeds.

#Download everything before losing network access

This is the decisive step and the only one that requires internet access. Do it comfortably at home or at the office over Wi-Fi, a few hours before departure—a model several GB in size doesn’t like unstable connections. Here’s how to do it with Ollama, the simplest approach for offline AI.

  1. 01
    Install Ollama while you have network access
    Download and install Ollama from ollama.com. This is the daemon that runs your models locally, listening on http://localhost:11434. Install it first, because the app itself is downloaded online.
  2. 02
    Download one or two compact models
    Use ollama pull to download the models to disk. Choose sizes suited to your machine (see the next section). A pull over stable Wi-Fi takes anywhere from a few minutes to a quarter of an hour, depending on the bandwidth and size.
  3. 03
    Install a graphical interface (optional)
    If you prefer a chat window to a terminal, install LM Studio (all-in-one, also handles downloads) or Open WebUI. LM Studio is particularly convenient on the go: app + models + interface in a single tool.
  4. 04
    Warm up each model once
    Start a test conversation with each downloaded model while you still have network access. This confirms that the file is complete and loads correctly into memory.
Terminal — preparation (with network access)
# Vérifier qu'Ollama est bien installé
ollama --version

# Télécharger des modèles compacts pour la mobilité
ollama pull qwen3.5:2b         # ~1,9 Go, très léger
ollama pull qwen3.5:9b         # ~6,6 Go, LE polyvalent 2026
ollama pull granite4.2:8b      # ~5,3 Go, sobre et token-efficient

# Lister ce qui est déjà sur le disque
ollama list
!
Don't start with an incomplete pull
An ollama pull interrompu leaves an incomplete model that will refuse to load offline. Wait for the “success” message and confirm the model is present with ollama list before cutting the network connection.

#Compact models that are sufficient for mobility

When working offline on a laptop, efficiency comes first: a small, fast model that uses little battery is better than a behemoth that crawls and drains the battery. For most mobile use cases—writing, summarization, general questions, coding assistance—a 2B to 9B model in Q4_K_M does the job very well.

Qwen 3.5 2B
~1.9 GB. The lightest option: ideal for an old laptop or for preserving battery life. Multimodal, 256k context, Apache 2.0 license. Sufficient for drafting, rewriting, and answering simple questions.
Qwen 3.5 9B
~6.6 GB. Excellent all-rounder, strong at reasoning, vision, and coding, with 256k context. The best compromise for a 16 GB RAM laptop—the reference 8 GB choice in 2026.
Granite 4.2 8B
~5.3 GB. Lightweight and token-efficient (IBM, Apache 2.0), 128k context, concise responses. A reliable choice for writing that preserves battery life.
Granite 4.2 3B
~2.2 GB. Compact and nimble, highly efficient, designed for modest machines while remaining capable.
→
Bring two models, not ten
A small one (2-3B) to save battery and respond quickly, and a medium one (8-9B) for demanding tasks. That's enough to cover 95% of offline needs without saturating the disk.

Q4_K_M quantization is the recommended default: it greatly reduces the model size and required memory, with minimal, imperceptible quality loss for most use cases. If you have RAM/VRAM to spare and want a little more precision, Q5_K_M is an alternative. Reserve Q8_0 and FP16 for very well-equipped machines.


#Verify that everything really works offline

Never take claims on faith: test in real-world conditions before you find yourself without a network connection. A three-minute test prevents disappointment at 10,000 feet.

  1. 01
    Cut off the network for good
    Turn on airplane mode, disable Wi-Fi AND Ethernet. The goal is to exactly reproduce a complete lack of connectivity.
  2. 02
    Start a conversation
    Open your terminal or LM Studio and chat with each downloaded model. If the model responds, your offline AI is operational.
  3. 03
    Check that no network error is blocking access
    Some interfaces display an offline warning banner: as long as generation works, ignore it. If the chat refuses to start, the model probably wasn’t fully downloaded.
Terminal — offline test (airplane mode enabled)
# Le daemon tourne en local, aucun réseau requis
ollama run qwen3.5:9b "Résume en trois points les avantages d'une IA hors ligne."

# Vérifier que le serveur local répond bien sur son port par défaut
curl http://localhost:11434/api/tags
i
The server remains local
Ollama listens on http://localhost:11434, an address internal to your machine. It has nothing to do with the internet: it works identically with or without a connection.

#Isolated workstation and air-gapped use

Some machines are deliberately kept permanently offline: workstations processing confidential data, industrial environments, and laboratories. This is the “air-gapped” use case—the target machine will never have network access, including for installation.

The principle: prepare the files on a connected machine, then transfer them via USB drive or external disk to the isolated workstation. Ollama stores its models in a directory that can be copied as is.

Offline installers
Retrieve the installation binary for Ollama or LM Studio from a connected machine, then copy it to the isolated workstation.
Model files
On Linux/macOS, Ollama models live in ~/.ollama/models; on Windows, in %USERPROFILE%\.ollama\models. Copy this folder to the target machine in the same location.
GGUF Alternative
With LM Studio, simply copy the downloaded .gguf files into the app’s model folder on the isolated machine.
Terminal — transfer to an isolated machine
# Sur la machine connectée : localiser les modèles
ls ~/.ollama/models

# Copier tout le dossier .ollama vers une clé USB
cp -r ~/.ollama /media/usb/ollama-backup

# Sur le poste air-gapped : restaurer au même emplacement
cp -r /media/usb/ollama-backup ~/.ollama

# Vérifier que les modèles sont bien reconnus
ollama list
!
Full copy required
The models folder contains both blobs (the weights) and manifests (the metadata). Copy the entire folder: a manifest without its blob, or vice versa, makes the model unusable on the target machine.

#Update your models when the network comes back

Once you're back in a covered area, it's time to maintain your setup: download new model versions, add new models you discovered along the way, and update the application. None of this can be done offline, which is why it's important to do it as soon as the network returns.

  1. 01
    Refresh existing models
    A ollama pull sur model that is already present downloads only the modified layers when a new version is available. It's fast and keeps your models up to date.
  2. 02
    Update the application
    Rerun the Ollama or LM Studio installer to get bug fixes and support for new models. Updates are frequent in this ecosystem.
  3. 03
    Clean up
    Delete models you no longer use to free disk space before your next trip.
  4. 04
    Re-test in airplane mode
    After every update, run an offline test: a new app version could introduce unexpected behavior without network access.
Terminal — maintenance (with network)
# Mettre à jour un modèle vers sa dernière version
ollama pull qwen3.5:9b

# Supprimer un modèle devenu inutile
ollama rm qwen3.5:2b

# Voir l'espace occupé par chaque modèle
ollama list

#Troubleshooting

“Model not found” offline
The model was not (completely) downloaded before the interruption. Reconnect, rerun ollama pull jusqu'until the success message, then verify with ollama list.
The app refuses to start without a network connection
Some interfaces attempt an online check at startup. Wait a few seconds for it to time out, or disable update checks in the settings.
Very slow responses while traveling
On pure CPU and battery power, this is normal. Switch to a smaller model (3B) and close resource-intensive applications. Plug in the laptop if possible.
Battery draining before your eyes
Inference puts a heavy load on the processor. Reduce the model size, limit response length, and enable power-saving mode for non-urgent tasks.
Copied model that won't load (air gap)
The copy was incomplete: blobs or manifests were missing. Copy the entire .ollama/models folder again.
Not enough memory
The model exceeds your RAM/VRAM. Choose a smaller size or a lighter quantization (Q4_K_M rather than Q8_0).

#Go further

Offline AI relies on the same foundations as a conventional local installation. These site guides complete the preparation:

Install Ollama
The step to perform while you still have network access to set up the inference engine.
Run an LLM locally without a GPU (CPU only)
Essential for mobility: expected models and speeds when running on battery in pure CPU mode.
Choose your quantization (Q4, Q5, Q8, FP16)
To balance disk size, memory, and quality based on the machine you carry with you.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.