AirLLM: a 70B on a 4 GB GPU, really? Test and limites
Air LLM makes a spectacular promise: running a 70-billion-parameter model on a GPU with only 4 GB of VRAM, when the rule of thumb says it needs around forty. The promise is real—but it comes at a price, and that price is speed. This guide explains how layer streaming works, tests the actual tokens/s achieved, and gives an honest verdict on when AirLLM is useful and when it is mainly a publicity stunt.
#AirLLM's promise: a 70B on 4 GB of VRAM
A 70B model in Q4 quantization requires about 40 GB of VRAM to be loaded entirely into memory—that is, one RTX 4090 (24 GB) plus a second card, or a Mac Studio with a large amount of unified memory. AirLLM claims to fit the same model on a 4 GB card, including a modest GTX 1650 or a free Google Colab T4. This is not a marketing trick involving the model size: it really is the full 70B model, in FP16 or 4/8-bit, producing the responses.
The secret comes down to three words: the model is never loaded entirely into VRAM. AirLLM splits the network into layers, stores them on disk, and loads only one layer into the GPU at a time during computation. VRAM therefore contains only the weights for one layer plus the current activations—hence the need for just a few gigabytes.
#How layer streaming works
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A transformer-based LLM is a stack of layers that are structurally identical (attention + feed-forward), traversed one after another. A Llama 70B has 80 of them. Standard inference loads all 80 layers into VRAM at once and keeps them resident in memory throughout generation. AirLLM reverses this principle.
- Disk partitioning
- On first load, AirLLM splits the model weights into separate files, one per layer, stored on the SSD. This conversion step happens only once per model.
- On-demand loading
- To generate a token, the engine loads layer 1 into the GPU, computes, frees the VRAM, loads layer 2, computes, and continues this way through layer 80.
- Constant VRAM
- At any given time, only one layer resides in VRAM. Peak memory usage depends on the largest layer and the context length, not the model’s total size.
- Prefetch and disk quantization
- AirLLM preloads the next layer while calculating the current layer and can compress weights on disk (block-wise quantization) to reduce the amount of data that must be read.
The bottleneck is obvious: every generated token requires reading the model's entire set of weights from disk. For a 70B model, that means transferring several dozen gigabytes from the SSD to the GPU for a single token. That's where all of AirLLM's performance—and all of its limitations—come from.
#Requirements and installation
AirLLM is a Python library built on PyTorch and the Hugging Face ecosystem. It isn't used like Ollama (no daemon, no run command): you import it into a Python script. Here's what you need before getting started.
- GPU
- Any NVIDIA card with at least 4 GB of VRAM (CUDA). Apple Silicon support (MPS) and CPU support exist but are even slower.
- Disk
- A fast NVMe SSD is essential, along with enough space to store the decompressed model (≈40 GB for a 70B, more in FP16).
- System RAM
- 16 GB is enough; AirLLM does not load the model into RAM, unlike a standard CPU offload.
- Python
- A Python 3.10+ environment with PyTorch and CUDA correctly installed.
#Running a 70B on 4 GB: the test
The minimal code to load and query a Llama 70B model fits in about fifteen lines. AirLLM handles splitting it into layers on the first call (allow several minutes for conversion and downloading the model from Hugging Face).
- 011. First loadAirLLM downloads the model and then converts it into per-layer files on the SSD. This step is long but one-time: subsequent runs reuse the disk cache.
- 022. Compression settingsThe compression='4bit' (or '8bit') parameter reduces the amount read from disk and therefore speeds up inference, at the cost of a slight loss in quality—the same trade-off as standard quantization.
- 033. GenerationEach token triggers a complete read of all 80 layers from the SSD. The progress bar advances layer by layer: it looks slow, and that's normal.
- 044. MeasurementTime the total duration and divide it by the number of generated tokens to get your actual token throughput in tokens/s. This is the only number that matters when deciding whether the tool is usable in your case.
#Real-world benchmark: how many tokens per second?
This is where the promise meets physical reality. With a 70B model streamed from an NVMe SSD, you're not talking about tokens per second but often seconds per token. The order of magnitude to keep in mind, based on reported configurations and our tests:
- 70B / fast PCIe 4.0 NVMe
- on the order of 0.1 to 0.5 tokens/s, or 2 to 10 seconds to produce a single word. A 200-token response takes several minutes.
- 70B / SATA SSD
- Another 2 to 5x slower: disk bandwidth tops out at ~500 MB/s, dropping below 0.1 token/s.
- Comparison of 70B loaded in VRAM (2x RTX 4090)
- 15 to 30 tokens/s. The gap with AirLLM is a factor of 30 to 300 depending on the drive.
- 8B model in Q4 on a single RTX 3060 12 GB
- 40 to 80 tokens/s, with no streaming at all — to illustrate the scale of the lost comfort.
The formula is simple: for every token, the entire model must be reread from disk. A 70B in 4-bit weighs ~40 GB; at 5 GB/s of NVMe read speed, that already means 8 seconds of pure I/O per token, before any computation. No software optimization can overcome this barrier as long as the weights live on disk. In the best case, this is a 5x to 30x slowdown (small models, very fast disk, aggressive compression), and much more for large models.
#Honest use cases
At 0.2 tokens/s, AirLLM is unusable for chatting. But there are real scenarios where the slowness is not a problem because no one is waiting in front of the screen.
- Offline batch processing
- A script that has to process 500 documents through a 70B overnight does not care if one document takes 3 minutes: the result is ready in the morning. This is the most legitimate use case.
- One-off experimentation
- Check what a specific 70B model answers for a few prompts, without renting a cloud GPU or buying hardware, to decide whether it is worth the investment.
- Non-urgent structured extraction
- Generate a dataset, annotate a corpus, produce embeddings or summaries as background tasks, where throughput matters little.
- Access to a giant model without a budget
- A student, researcher, or curious user with just one small card who wants to try a model they otherwise could not run.
#When it's marketing
The claim “a 70B model on 4 GB” is technically true but editorially misleading whenever it implies normal use. Here are the situations where AirLLM falls short of its implicit promises.
- Interactive chat
- Waiting several minutes for a response kills any conversation. For discussion, an instant local 8B or 14B is infinitely more useful than a 70B that responds at fax speed.
- Real-time coding assistant
- Autocomplete and pair programming require responses within a few seconds. AirLLM is light-years away from that.
- Multi-user server
- Impossible to serve multiple people: each token already monopolizes the entire disk bandwidth for a single request.
- Production
- No online service can rely on AirLLM. Latency and SSD wear (constant massive reads) rule it out from the start.
#Alternatives: start with standard quantization
Before resorting to layer streaming, exhaust the options that keep the model in memory—they are almost always preferable. The question isn't “how do I fit a 70B model into 4 GB?” but “which model actually meets my needs?”
- Lowering the quantization
- A 70B in Q4_K_M fits in ~40 GB; in Q2/Q3, it takes much less. But an overly aggressively quantized 70B loses quality: a well-loaded 32B in Q4 is often a better choice.
- Choose a smaller, newer model
- A 32B (≈19 GB VRAM) or a 14B (≈9 GB) from 2026 rivals a 2024 70B on most tasks, while running instantly on a single card.
- CPU/GPU offload for Ollama or llama.cpp
- These engines offload some layers to system RAM when VRAM is insufficient. It’s slower than a full-GPU load, but much faster than AirLLM because RAM is a hundred times faster than an SSD.
- Renting a cloud GPU for an occasional need
- To test a real 70B quickly just once, an hour of rented GPU time costs little and provides normal throughput—often more cost-effective than waiting hours locally.
#Go further
AirLLM is a workaround tool: before turning to it, it is better to master the standard VRAM and model-selection levers. Three guides on the site round out the picture.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand how much quality a model loses in Q4 or Q3, and how to fit a large model into limited VRAM without resorting to disk streaming.
- Choose your GPU for local AI
- The VRAM/budget/performance trade-offs, from the RTX 3060 12 GB to the Mac Studio Ultra, so you can determine what model size your hardware can really handle.
- llama.cpp vs vLLM vs Exllama
- Inference engines and their CPU/GPU offloading support—the serious alternative to AirLLM when VRAM is slightly insufficient.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.