Intermediate 10 minInference

AirLLM: a 70B on a 4 GB GPU, really? Test and limites

Air LLM makes a spectacular promise: running a 70-billion-parameter model on a GPU with only 4 GB of VRAM, when the rule of thumb says it needs around forty. The promise is real—but it comes at a price, and that price is speed. This guide explains how layer streaming works, tests the actual tokens/s achieved, and gives an honest verdict on when AirLLM is useful and when it is mainly a publicity stunt.

By Mohamed Meguedmi·Update 2026-08-26·Tested on Windows, macOS, and Linux

#AirLLM's promise: a 70B on 4 GB of VRAM

A 70B model in Q4 quantization requires about 40 GB of VRAM to be loaded entirely into memory—that is, one RTX 4090 (24 GB) plus a second card, or a Mac Studio with a large amount of unified memory. AirLLM claims to fit the same model on a 4 GB card, including a modest GTX 1650 or a free Google Colab T4. This is not a marketing trick involving the model size: it really is the full 70B model, in FP16 or 4/8-bit, producing the responses.

The secret comes down to three words: the model is never loaded entirely into VRAM. AirLLM splits the network into layers, stores them on disk, and loads only one layer into the GPU at a time during computation. VRAM therefore contains only the weights for one layer plus the current activations—hence the need for just a few gigabytes.

i
What this doesn’t change
AirLLM does not compress the model beyond the selected quantization, nor does it degrade its response quality: the computation is mathematically identical to that of a model loaded in full. The only thing that changes is where the weights reside between computations—on disk rather than in VRAM.

#How layer streaming works

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A transformer-based LLM is a stack of layers that are structurally identical (attention + feed-forward), traversed one after another. A Llama 70B has 80 of them. Standard inference loads all 80 layers into VRAM at once and keeps them resident in memory throughout generation. AirLLM reverses this principle.

Disk partitioning
On first load, AirLLM splits the model weights into separate files, one per layer, stored on the SSD. This conversion step happens only once per model.
On-demand loading
To generate a token, the engine loads layer 1 into the GPU, computes, frees the VRAM, loads layer 2, computes, and continues this way through layer 80.
Constant VRAM
At any given time, only one layer resides in VRAM. Peak memory usage depends on the largest layer and the context length, not the model’s total size.
Prefetch and disk quantization
AirLLM preloads the next layer while calculating the current layer and can compress weights on disk (block-wise quantization) to reduce the amount of data that must be read.

The bottleneck is obvious: every generated token requires reading the model's entire set of weights from disk. For a 70B model, that means transferring several dozen gigabytes from the SSD to the GPU for a single token. That's where all of AirLLM's performance—and all of its limitations—come from.

!
The disk becomes the real processor
With layer streaming, GPU compute power is no longer what limits speed; your SSD's bandwidth is. An NVMe PCIe 4.0 (~5 GB/s) and an old SATA SSD (~500 MB/s) produce radically different results. A mechanical hard drive is simply unusable.

#Requirements and installation

AirLLM is a Python library built on PyTorch and the Hugging Face ecosystem. It isn't used like Ollama (no daemon, no run command): you import it into a Python script. Here's what you need before getting started.

GPU
Any NVIDIA card with at least 4 GB of VRAM (CUDA). Apple Silicon support (MPS) and CPU support exist but are even slower.
Disk
A fast NVMe SSD is essential, along with enough space to store the decompressed model (≈40 GB for a 70B, more in FP16).
System RAM
16 GB is enough; AirLLM does not load the model into RAM, unlike a standard CPU offload.
Python
A Python 3.10+ environment with PyTorch and CUDA correctly installed.
Installation
pip install airllm
i
This is not an alternative to Ollama
AirLLM does not replace Ollama or LM Studio for daily use. It is a niche tool geared toward Python scripting, best reserved for situations where you truly lack the VRAM for a model and slowness is not a deal-breaker.

#Running a 70B on 4 GB: the test

The minimal code to load and query a Llama 70B model fits in about fifteen lines. AirLLM handles splitting it into layers on the first call (allow several minutes for conversion and downloading the model from Hugging Face).

70B inference with AirLLM
from airllm import AutoModel

# Le modèle est découpé en couches au premier chargement
model = AutoModel.from_pretrained("meta-llama/Meta-Llama-3.1-70B-Instruct")

input_text = ["Explique le layer streaming en une phrase."]
input_tokens = model.tokenizer(input_text,
                               return_tensors="pt",
                               truncation=True,
                               max_length=128,
                               padding=False)

generation_output = model.generate(
    input_tokens['input_ids'].cuda(),
    max_new_tokens=64,
    use_cache=True,
    return_dict_in_generate=True)

output = model.tokenizer.decode(generation_output.sequences[0])
print(output)
  1. 01
    1. First load
    AirLLM downloads the model and then converts it into per-layer files on the SSD. This step is long but one-time: subsequent runs reuse the disk cache.
  2. 02
    2. Compression settings
    The compression='4bit' (or '8bit') parameter reduces the amount read from disk and therefore speeds up inference, at the cost of a slight loss in quality—the same trade-off as standard quantization.
  3. 03
    3. Generation
    Each token triggers a complete read of all 80 layers from the SSD. The progress bar advances layer by layer: it looks slow, and that's normal.
  4. 04
    4. Measurement
    Time the total duration and divide it by the number of generated tokens to get your actual token throughput in tokens/s. This is the only number that matters when deciding whether the tool is usable in your case.

#Real-world benchmark: how many tokens per second?

This is where the promise meets physical reality. With a 70B model streamed from an NVMe SSD, you're not talking about tokens per second but often seconds per token. The order of magnitude to keep in mind, based on reported configurations and our tests:

70B / fast PCIe 4.0 NVMe
on the order of 0.1 to 0.5 tokens/s, or 2 to 10 seconds to produce a single word. A 200-token response takes several minutes.
70B / SATA SSD
Another 2 to 5x slower: disk bandwidth tops out at ~500 MB/s, dropping below 0.1 token/s.
Comparison of 70B loaded in VRAM (2x RTX 4090)
15 to 30 tokens/s. The gap with AirLLM is a factor of 30 to 300 depending on the drive.
8B model in Q4 on a single RTX 3060 12 GB
40 to 80 tokens/s, with no streaming at all — to illustrate the scale of the lost comfort.

The formula is simple: for every token, the entire model must be reread from disk. A 70B in 4-bit weighs ~40 GB; at 5 GB/s of NVMe read speed, that already means 8 seconds of pure I/O per token, before any computation. No software optimization can overcome this barrier as long as the weights live on disk. In the best case, this is a 5x to 30x slowdown (small models, very fast disk, aggressive compression), and much more for large models.

→
Prefetch helps a little
AirLLM preloads the next layer while calculating the current layer, masking some of the disk latency. That is what distinguishes AirLLM from simple naive offloading—but it does not change the order of magnitude: throughput is still dictated by SSD speed.

#Honest use cases

At 0.2 tokens/s, AirLLM is unusable for chatting. But there are real scenarios where the slowness is not a problem because no one is waiting in front of the screen.

Offline batch processing
A script that has to process 500 documents through a 70B overnight does not care if one document takes 3 minutes: the result is ready in the morning. This is the most legitimate use case.
One-off experimentation
Check what a specific 70B model answers for a few prompts, without renting a cloud GPU or buying hardware, to decide whether it is worth the investment.
Non-urgent structured extraction
Generate a dataset, annotate a corpus, produce embeddings or summaries as background tasks, where throughput matters little.
Access to a giant model without a budget
A student, researcher, or curious user with just one small card who wants to try a model they otherwise could not run.

#When it's marketing

The claim “a 70B model on 4 GB” is technically true but editorially misleading whenever it implies normal use. Here are the situations where AirLLM falls short of its implicit promises.

Interactive chat
Waiting several minutes for a response kills any conversation. For discussion, an instant local 8B or 14B is infinitely more useful than a 70B that responds at fax speed.
Real-time coding assistant
Autocomplete and pair programming require responses within a few seconds. AirLLM is light-years away from that.
Multi-user server
Impossible to serve multiple people: each token already monopolizes the entire disk bandwidth for a single request.
Production
No online service can rely on AirLLM. Latency and SSD wear (constant massive reads) rule it out from the start.
!
SSD wear is not trivial
Streaming a 70B continuously means reading tens of GB per token, potentially amounting to terabytes read in a single session. Reads do not wear out NAND cells like writes do, but the initial conversion and cache also write a lot. Reserve this for occasional use, not a 24/7 loop.

#Alternatives: start with standard quantization

Before resorting to layer streaming, exhaust the options that keep the model in memory—they are almost always preferable. The question isn't “how do I fit a 70B model into 4 GB?” but “which model actually meets my needs?”

Lowering the quantization
A 70B in Q4_K_M fits in ~40 GB; in Q2/Q3, it takes much less. But an overly aggressively quantized 70B loses quality: a well-loaded 32B in Q4 is often a better choice.
Choose a smaller, newer model
A 32B (≈19 GB VRAM) or a 14B (≈9 GB) from 2026 rivals a 2024 70B on most tasks, while running instantly on a single card.
CPU/GPU offload for Ollama or llama.cpp
These engines offload some layers to system RAM when VRAM is insufficient. It’s slower than a full-GPU load, but much faster than AirLLM because RAM is a hundred times faster than an SSD.
Renting a cloud GPU for an occasional need
To test a real 70B quickly just once, an hour of rented GPU time costs little and provides normal throughput—often more cost-effective than waiting hours locally.
→
The right question
In 90% of cases, the real answer to “I don’t have enough VRAM for a 70B” is “get a 32B in Q4.” You will get similar quality, 100x the throughput, and zero disk streaming. AirLLM is justified only if the exact 70B model is non-negotiable AND the slowness is acceptable.

#Go further

AirLLM is a workaround tool: before turning to it, it is better to master the standard VRAM and model-selection levers. Three guides on the site round out the picture.

Choose your quantization (Q4, Q5, Q8, FP16)
Understand how much quality a model loses in Q4 or Q3, and how to fit a large model into limited VRAM without resorting to disk streaming.
Choose your GPU for local AI
The VRAM/budget/performance trade-offs, from the RTX 3060 12 GB to the Mac Studio Ultra, so you can determine what model size your hardware can really handle.
llama.cpp vs vLLM vs Exllama
Inference engines and their CPU/GPU offloading support—the serious alternative to AirLLM when VRAM is slightly insufficient.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.