Liquid AI's LFM2: the alternative architecture for l'edge
Liquid AI's LFM2 is not just another transformer: it is a family of small models (350M to 8B) built on a hybrid architecture in which most layers are short convolutions rather than attention. The result announced by Liquid AI: decoding and prefill are about twice as fast as Qwen3 at the same size on CPU, with memory usage growing little as context increases. This guide explains what this architecture really changes, how to install LFM2 with Ollama or llama.cpp, what speed to expect without a GPU, and which use cases this choice beats a conventional transformer for.
#Liquid AI's LFM2: why a model designed for the CPU
Liquid AI is a company spun out of MIT (CSAIL), founded in 2023 around “liquid neural networks” and continuous-state models. After an initial closed LFM1 generation, the company released the LFM2 weights in July 2025, its second generation, in three sizes: 350M, 700M, and 1.2B parameters. Other variants followed (2.6B, an 8B-A1B MoE version, vision and audio models, and then the LFM2.5 generation). The through line remains the same: these models target on-device execution—that is, on a phone, a laptop without a graphics card, a mini-PC, or an embedded board.
Almost all the small open models you know (Qwen3, Gemma 3, Llama 3.2, SmolLM) are conventional transformers: each layer applies attention across the entire context. This works very well on a GPU, but on a CPU each generated token has to reread a key-value cache (KV cache) that grows with the conversation, and prefilling a long prompt is expensive. LFM2 targets precisely these two issues by replacing most attention layers with short-convolution blocks that use far less memory and bandwidth.
- Target
- On-device inference: x86 or ARM CPU, NPU, integrated GPU. The dedicated GPU is not the primary playing field, even if it works.
- Quantified promise
- Liquid AI claims decoding and prefill up to 2 times faster than Qwen3 at a comparable size on CPU (measurements published on an AMD Ryzen AI 9 HX 370 and a Samsung Galaxy S24 Ultra).
- Quality
- With the same parameter count, LFM2 performs at or slightly above competing transformers on knowledge, instruction, and math benchmarks; the 1.2B rivals Qwen3-1.7B on MMLU and IFEval.
- License
- LFM Open License v1.0: unrestricted commercial use below an annual revenue threshold ($10 million); above that, you must contact Liquid AI. This is not Apache 2.0, so reread the text before an enterprise deployment.
#What the LFM hybrid architecture changes
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
LFM2 stacks 16 blocks. Ten are double-gated short-range convolution blocks, and six are grouped query attention blocks (GQA), as in a modern transformer. Liquid AI describes these convolution blocks as LIV (linear input-varying) operators: the weights applied at each position depend on the input, giving the block a form of selectivity similar to state-space models (Mamba) or modern RNNs, without their complex recurrent state.
In practice, a convolution block looks at only a few neighboring tokens. Its per-token cost is constant, regardless of the context already generated, and it has nothing to store in a KV cache. Only the six attention blocks retain a cache, compared with 28 or 36 layers in a similarly sized transformer. Two direct consequences on CPU:
- Faster decoding
- Generating a token mainly means rereading the weights and KV cache from RAM. With less cache to reread, memory bandwidth—the true bottleneck of a CPU—is used more effectively.
- Efficient prefill
- Ingesting a 4,000-token prompt (a document for RAG, a chat history) costs proportionally less than with a full transformer because ten of sixteen blocks work in a local window.
- Stable memory
- Consumption increases slowly with context: the KV cache for six layers remains small, which matters on a phone or a card with 4 or 8 GB of shared RAM.
- Native 32k context
- LFM2 was trained with a 32,768-token context window (128k on some newer variants), enough for summarization and lightweight RAG.
For training, Liquid AI used about 10,000 billion tokens for the first-generation LFM2, with a mix dominated by English, about 20% multilingual data (including French, German, Spanish, Arabic, Chinese, Japanese, and Korean), and some code, plus distillation from the internal LFM1-7B. That's why an LFM2-1.2B can hold a decent conversation in French, which cannot be taken for granted for all models of this size.
#The LFM2 family: sizes and variants
All variants share the same base architecture and chat format (im_start/im_end tags, in the ChatML style). Here are the ones that matter for edge use, with the approximate GGUF file size in Q4_K_M, which roughly corresponds to the RAM occupied by the weights on CPU.
- LFM2-350M
- About 250 MB in Q4. Classification, extraction, short rewriting. Runs on almost anything, including a Raspberry Pi or an old laptop.
- LFM2-700M
- About 450 MB in Q4. The right compromise for a recent phone or a very lightweight assistant.
- LFM2-1.2B
- About 730 MB in Q4. The family's reference model: chat, summarization, basic RAG, and tool calls. This is the one the guide installs.
- LFM2-2.6B
- About 1.5 GB in Q4. Released in late 2025, it's significantly stronger at reasoning and multilingual tasks, while still running comfortably on a laptop without a GPU.
- LFM2-8B-A1B
- Mixture of Experts: 8.3B total parameters, with around 1.5B active per token. Around 5 GB in Q4: it needs the RAM of an 8B model, but its speed remains that of a small model.
- Specialized variants
- LFM2-VL (vision, 450M and 1.6B), LFM2-Audio, and fine-tuned variants for data extraction, RAG, or tool calls. The LFM2.5 generation (starting in January 2026) retains the architecture with extended training; the site's catalog lists its model pages, including the 2.6B and 7B versions.
#Prerequisites
Nothing exotic. The important point is to use a recent version of the inference engine: support for the LFM2 architecture was added to llama.cpp in July 2025 and to Hugging Face Transformers in version 4.54. Versions of Ollama and LM Studio released since then include this support, provided you update them.
- Machine
- Any recent PC or Mac. A 4-core CPU and 8 GB of RAM are enough for the 1.2B; allow 16 GB for the 8B-A1B.
- Ollama up to date
- Ollama listens on http://localhost:11434 by default. Update it before pulling the model: any version from before summer 2025 will reject the GGUF with an unknown architecture error.
- Or llama.cpp
- A recent binary (compiled or downloaded from the GitHub releases) provides access to llama-cli, llama-server, and especially llama-bench for measuring speed.
- Python optional
- To use the original (unquantized) weights with Transformers ≥ 4.54, for example for fine-tuning or exporting to a mobile SDK.
#Install and test LFM2 locally
Liquid AI publishes its models on Hugging Face under the LiquidAI organization, with an original-weights repository for each size (LiquidAI/LFM2-1.2B) and an already-quantized GGUF repository (LiquidAI/LFM2-1.2B-GGUF). The simplest approach is to pull this GGUF directly into Ollama, without going through a Modelfile.
- 01Update OllamaOn Linux, rerun the official installation script; on macOS and Windows, the app updates itself or through its menu. Check with ollama --version.
- 02Pull the GGUF from Hugging FaceThe hf.co/organisation/dépôt:quantification syntax works with any public GGUF repository. For LFM2-1.2B in Q4_K_M, the download is approximately 730 MB.
- 03Start a first chatollama run ouvre une session interactive. Posez une question en français pour vérifier la qualité de la langue avant de creuser.
- 04Force CPU mode for comparisonIf your machine has a GPU, you can disable it for this model by setting the num_gpu parameter to 0 and observe its behavior on pure CPU.
- 05Connect an interfaceThe model immediately appears in Open WebUI, LM Studio, or any OpenAI-compatible client pointed at port 11434.
The name hf.co/LiquidAI/LFM2-1.2B-GGUF:Q4_K_M is long to type; create a local alias with ollama cp to get a simple lfm2:1.2b. For the other sizes, replace 1.2B with 350M, 700M, or 2.6B, and for the MoE use the LiquidAI/LFM2-8B-A1B-GGUF repository.
If you prefer llama.cpp directly, the llama-cli command accepts the same Hugging Face repository with the -hf option. It is also the preferred route for a minimalist server on an ARM board, where llama-server runs with a smaller footprint than Ollama.
Finally, for a Python script using the original bfloat16 weights, Transformers is enough. Allow about 2.5 GB of RAM for the 1.2B model in bf16; on CPU, this path is much slower than llama.cpp and is only useful for development.
#Pure CPU speed: the real numbers
On a CPU, two numbers matter: prefill speed (prompt tokens processed per second, which determines the time before the first word) and generation speed (tokens produced per second). The right tool for measuring them properly is llama-bench, shipped with llama.cpp: it isolates the two phases and repeats the measurements. Force the CPU with -ngl 0 even if a GPU is present.
The orders of magnitude below correspond to what we observe with llama.cpp in Q4_K_M, using all physical cores, on typical machines. They do not replace your own measurements: memory bandwidth (DDR4 versus DDR5, number of channels) can make the result vary twofold between two PCs of the same generation.
- Recent laptop (8 cores, DDR5)
- LFM2-1.2B: 40 to 70 tokens/s during generation, and several hundred tokens/s during prefill. The 350M exceeds 100 tokens/s. The 2.6B runs at around 25 to 40 tokens/s.
- Desktop PC with 4 to 6 cores, DDR4
- LFM2-1.2B: 20 to 35 tokens/s, well above reading speed. A Qwen3-1.7B on the same machine is closer to 12 to 20 tokens/s.
- Mac Apple Silicon (CPU only)
- Comparable to a DDR5 laptop thanks to unified memory; in practice, you’ll let Metal accelerate it, but LFM2 remains comfortable even on CPU alone on a MacBook Air.
- Raspberry Pi 5 (8 GB)
- LFM2-1.2B: 8 to 12 tokens/s; LFM2-350M: 25 to 35 tokens/s. This is where the gap with a classic transformer is most visible.
- LFM2-8B-A1B
- With 16 GB of DDR5 RAM, expect 30 to 50 tokens/s: only 1.5B parameters work on each token, but the 5 GB of weights must fit in memory.
The factor that makes the difference in practice is prefilling. On a 1.7B transformer, ingesting a 4,000-token document on the CPU often takes 15 to 30 seconds before the first response; LFM2-1.2B cuts that time by nearly half in tests published by Liquid AI, making local RAG on the CPU genuinely usable rather than merely possible.
#Which use cases favor LFM2
LFM2 is not a replacement for Qwen3-8B or Gemma 3 12B. On a GPU with VRAM, a larger transformer will be smarter, period. LFM2 becomes useful as soon as hardware is the constraint: no GPU, limited RAM, a battery to conserve, or a query volume to handle at a fixed cost.
- Laptop assistant without a GPU
- A smooth chat on a fanless ultraportable, without the fan spinning up. The 1.2B or 2.6B respond faster than you can read.
- Batch processing on a CPU server
- Classify tickets, extract fields, rewrite product descriptions: thousands of short queries per hour on a VM without a GPU, using the 350M or 700M.
- Lightweight local RAG
- Fast prefill + 32k context: index notes or internal documentation and answer on CPU in a few seconds.
- Embedded and home automation
- Raspberry Pi, industrial mini-PC, home automation hub: interpret a transcribed voice command, generate a short response, call a tool.
- Mobile
- Liquid AI offers its LEAP SDK (Liquid Edge AI Platform) for iOS and Android, along with the Apollo app for testing models on a phone. GGUF files also work in apps based on llama.cpp.
- Tool calls
- LFM2 instruct variants natively support function calling with dedicated tags, making them an affordable small agent router.
Conversely, stick with a conventional transformer when you have a GPU and VRAM to fill, when the task requires long-form reasoning (math, complex code), or when you depend on a well-stocked fine-tuning ecosystem: Qwen and Llama still have the edge in this area.
#Limitations and troubleshooting
- “unknown model architecture: lfm2” error
- Your engine is too old. Update Ollama, LM Studio, or recompile llama.cpp from a version released after July 2025.
- Responses that get stuck in a loop
- Lower the temperature to 0.3 and enable min_p 0.15 and repeat_penalty 1.05, the values recommended by Liquid AI. Small models are sensitive to the default settings.
- Disappointing speed
- Check with ollama ps that the model is fully loaded, reduce the thread count to the performance cores, and close applications that saturate memory bandwidth (a browser with dozens of tabs, for example).
- Bad French
- The 350M remains limited in French; move up to the 1.2B or 2.6B, trained with a more useful multilingual component.
- Overly aggressive quantization
- With a 350M or 700M model, Q4 damages quality more than it does on a 7B model. Prefer Q8_0 for smaller sizes: the file remains tiny (under 800 MB for the 700M).
- Enterprise licensing
- Beyond the revenue threshold set by the LFM Open License, commercial use requires an agreement with Liquid AI. Check before putting it into production.
#Go further
LFM2 makes the most sense in a setup without a GPU or on a small machine. These site guides complement this one:
- Local LLM without a GPU (CPU)
- Recommended models by RAM capacity and CPU tokens/s benchmarks, to position LFM2 against the alternatives.
- Import a GGUF model from Hugging Face into Ollama
- Everything about hf.co syntax, Modelfiles, and aliases, useful for locking in the recommended LFM2 parameters.
- LLM on Raspberry Pi 5
- The quintessential embedded use case, where LFM2’s speed makes the difference.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.