Intermediate 10 minEdge

Liquid AI's LFM2: the alternative architecture for l'edge

Liquid AI's LFM2 is not just another transformer: it is a family of small models (350M to 8B) built on a hybrid architecture in which most layers are short convolutions rather than attention. The result announced by Liquid AI: decoding and prefill are about twice as fast as Qwen3 at the same size on CPU, with memory usage growing little as context increases. This guide explains what this architecture really changes, how to install LFM2 with Ollama or llama.cpp, what speed to expect without a GPU, and which use cases this choice beats a conventional transformer for.

By Thomas P.·Update 2026-09-24·Tested on Windows, macOS, and Linux

#Liquid AI's LFM2: why a model designed for the CPU

Liquid AI is a company spun out of MIT (CSAIL), founded in 2023 around “liquid neural networks” and continuous-state models. After an initial closed LFM1 generation, the company released the LFM2 weights in July 2025, its second generation, in three sizes: 350M, 700M, and 1.2B parameters. Other variants followed (2.6B, an 8B-A1B MoE version, vision and audio models, and then the LFM2.5 generation). The through line remains the same: these models target on-device execution—that is, on a phone, a laptop without a graphics card, a mini-PC, or an embedded board.

Almost all the small open models you know (Qwen3, Gemma 3, Llama 3.2, SmolLM) are conventional transformers: each layer applies attention across the entire context. This works very well on a GPU, but on a CPU each generated token has to reread a key-value cache (KV cache) that grows with the conversation, and prefilling a long prompt is expensive. LFM2 targets precisely these two issues by replacing most attention layers with short-convolution blocks that use far less memory and bandwidth.

Target
On-device inference: x86 or ARM CPU, NPU, integrated GPU. The dedicated GPU is not the primary playing field, even if it works.
Quantified promise
Liquid AI claims decoding and prefill up to 2 times faster than Qwen3 at a comparable size on CPU (measurements published on an AMD Ryzen AI 9 HX 370 and a Samsung Galaxy S24 Ultra).
Quality
With the same parameter count, LFM2 performs at or slightly above competing transformers on knowledge, instruction, and math benchmarks; the 1.2B rivals Qwen3-1.7B on MMLU and IFEval.
License
LFM Open License v1.0: unrestricted commercial use below an annual revenue threshold ($10 million); above that, you must contact Liquid AI. This is not Apache 2.0, so reread the text before an enterprise deployment.

#What the LFM hybrid architecture changes

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

LFM2 stacks 16 blocks. Ten are double-gated short-range convolution blocks, and six are grouped query attention blocks (GQA), as in a modern transformer. Liquid AI describes these convolution blocks as LIV (linear input-varying) operators: the weights applied at each position depend on the input, giving the block a form of selectivity similar to state-space models (Mamba) or modern RNNs, without their complex recurrent state.

In practice, a convolution block looks at only a few neighboring tokens. Its per-token cost is constant, regardless of the context already generated, and it has nothing to store in a KV cache. Only the six attention blocks retain a cache, compared with 28 or 36 layers in a similarly sized transformer. Two direct consequences on CPU:

Faster decoding
Generating a token mainly means rereading the weights and KV cache from RAM. With less cache to reread, memory bandwidth—the true bottleneck of a CPU—is used more effectively.
Efficient prefill
Ingesting a 4,000-token prompt (a document for RAG, a chat history) costs proportionally less than with a full transformer because ten of sixteen blocks work in a local window.
Stable memory
Consumption increases slowly with context: the KV cache for six layers remains small, which matters on a phone or a card with 4 or 8 GB of shared RAM.
Native 32k context
LFM2 was trained with a 32,768-token context window (128k on some newer variants), enough for summarization and lightweight RAG.

For training, Liquid AI used about 10,000 billion tokens for the first-generation LFM2, with a mix dominated by English, about 20% multilingual data (including French, German, Spanish, Arabic, Chinese, Japanese, and Korean), and some code, plus distillation from the internal LFM1-7B. That's why an LFM2-1.2B can hold a decent conversation in French, which cannot be taken for granted for all models of this size.

i
Hybrid, not revolutionary
LFM2 does not abandon attention: the six GQA blocks retain the ability to connect two distant passages in the context. It is the same tradeoff as Jamba (Mamba + attention) or IBM's recent Granite 4 models, applied to much smaller sizes. The difference lies in the ratio and the convolution operator chosen.

#The LFM2 family: sizes and variants

All variants share the same base architecture and chat format (im_start/im_end tags, in the ChatML style). Here are the ones that matter for edge use, with the approximate GGUF file size in Q4_K_M, which roughly corresponds to the RAM occupied by the weights on CPU.

LFM2-350M
About 250 MB in Q4. Classification, extraction, short rewriting. Runs on almost anything, including a Raspberry Pi or an old laptop.
LFM2-700M
About 450 MB in Q4. The right compromise for a recent phone or a very lightweight assistant.
LFM2-1.2B
About 730 MB in Q4. The family's reference model: chat, summarization, basic RAG, and tool calls. This is the one the guide installs.
LFM2-2.6B
About 1.5 GB in Q4. Released in late 2025, it's significantly stronger at reasoning and multilingual tasks, while still running comfortably on a laptop without a GPU.
LFM2-8B-A1B
Mixture of Experts: 8.3B total parameters, with around 1.5B active per token. Around 5 GB in Q4: it needs the RAM of an 8B model, but its speed remains that of a small model.
Specialized variants
LFM2-VL (vision, 450M and 1.6B), LFM2-Audio, and fine-tuned variants for data extraction, RAG, or tool calls. The LFM2.5 generation (starting in January 2026) retains the architecture with extended training; the site's catalog lists its model pages, including the 2.6B and 7B versions.
→
Which model should you choose?
On a laptop or mini PC with 8 GB of RAM: start with the 1.2B, then move up to the 2.6B if the quality is insufficient. On 16 GB: the 8B-A1B is the most capable while remaining fast. On a phone or Raspberry Pi: 350M or 700M.

#Prerequisites

Nothing exotic. The important point is to use a recent version of the inference engine: support for the LFM2 architecture was added to llama.cpp in July 2025 and to Hugging Face Transformers in version 4.54. Versions of Ollama and LM Studio released since then include this support, provided you update them.

Machine
Any recent PC or Mac. A 4-core CPU and 8 GB of RAM are enough for the 1.2B; allow 16 GB for the 8B-A1B.
Ollama up to date
Ollama listens on http://localhost:11434 by default. Update it before pulling the model: any version from before summer 2025 will reject the GGUF with an unknown architecture error.
Or llama.cpp
A recent binary (compiled or downloaded from the GitHub releases) provides access to llama-cli, llama-server, and especially llama-bench for measuring speed.
Python optional
To use the original (unquantized) weights with Transformers ≥ 4.54, for example for fine-tuning or exporting to a mobile SDK.

#Install and test LFM2 locally

Liquid AI publishes its models on Hugging Face under the LiquidAI organization, with an original-weights repository for each size (LiquidAI/LFM2-1.2B) and an already-quantized GGUF repository (LiquidAI/LFM2-1.2B-GGUF). The simplest approach is to pull this GGUF directly into Ollama, without going through a Modelfile.

  1. 01
    Update Ollama
    On Linux, rerun the official installation script; on macOS and Windows, the app updates itself or through its menu. Check with ollama --version.
  2. 02
    Pull the GGUF from Hugging Face
    The hf.co/organisation/dépôt:quantification syntax works with any public GGUF repository. For LFM2-1.2B in Q4_K_M, the download is approximately 730 MB.
  3. 03
    Start a first chat
    ollama run ouvre une session interactive. Posez une question en français pour vérifier la qualité de la langue avant de creuser.
  4. 04
    Force CPU mode for comparison
    If your machine has a GPU, you can disable it for this model by setting the num_gpu parameter to 0 and observe its behavior on pure CPU.
  5. 05
    Connect an interface
    The model immediately appears in Open WebUI, LM Studio, or any OpenAI-compatible client pointed at port 11434.
Terminal — Ollama
# Vérifier la version (doit être récente, été 2025 ou plus)
ollama --version

# Tirer LFM2-1.2B quantifié en Q4_K_M depuis Hugging Face
ollama pull hf.co/LiquidAI/LFM2-1.2B-GGUF:Q4_K_M

# Premier chat
ollama run hf.co/LiquidAI/LFM2-1.2B-GGUF:Q4_K_M

# Même chose en CPU pur, même si un GPU est présent
ollama run hf.co/LiquidAI/LFM2-1.2B-GGUF:Q4_K_M
>>> /set parameter num_gpu 0
>>> Résume en trois phrases ce qu'est un modèle de langage.

The name hf.co/LiquidAI/LFM2-1.2B-GGUF:Q4_K_M is long to type; create a local alias with ollama cp to get a simple lfm2:1.2b. For the other sizes, replace 1.2B with 350M, 700M, or 2.6B, and for the MoE use the LiquidAI/LFM2-8B-A1B-GGUF repository.

Terminal — aliases and API
# Alias court
ollama cp hf.co/LiquidAI/LFM2-1.2B-GGUF:Q4_K_M lfm2:1.2b

# Appel API (compatible avec n'importe quel client Ollama)
curl http://localhost:11434/api/chat -d '{
  "model": "lfm2:1.2b",
  "messages": [{"role": "user", "content": "Explique la différence entre RAM et VRAM en deux phrases."}],
  "options": {"temperature": 0.3, "min_p": 0.15, "repeat_penalty": 1.05},
  "stream": false
}'
→
Liquid AI’s recommended settings
The official page recommends temperature 0.3, min_p 0.15, and repetition_penalty 1.05. With the default values in Ollama (temperature 0.8), a 1.2B model is more likely to get stuck in a loop or go off-topic. Set these three parameters in a Modelfile if you use LFM2 in production.

If you prefer llama.cpp directly, the llama-cli command accepts the same Hugging Face repository with the -hf option. It is also the preferred route for a minimalist server on an ARM board, where llama-server runs with a smaller footprint than Ollama.

Terminal — llama.cpp
# Chat interactif, téléchargement automatique du GGUF
llama-cli -hf LiquidAI/LFM2-1.2B-GGUF:Q4_K_M -cnv \
  --temp 0.3 --min-p 0.15 --repeat-penalty 1.05 -t 8

# Serveur API OpenAI-compatible sur le port 8080
llama-server -hf LiquidAI/LFM2-1.2B-GGUF:Q4_K_M -c 8192 -t 8

Finally, for a Python script using the original bfloat16 weights, Transformers is enough. Allow about 2.5 GB of RAM for the 1.2B model in bf16; on CPU, this path is much slower than llama.cpp and is only useful for development.

Python — Transformers ≥ 4.54
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "LiquidAI/LFM2-1.2B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16")

messages = [{"role": "user", "content": "Qu'est-ce qu'un modèle hybride convolution + attention ?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=200, do_sample=True, temperature=0.3, min_p=0.15, repetition_penalty=1.05)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

#Pure CPU speed: the real numbers

On a CPU, two numbers matter: prefill speed (prompt tokens processed per second, which determines the time before the first word) and generation speed (tokens produced per second). The right tool for measuring them properly is llama-bench, shipped with llama.cpp: it isolates the two phases and repeats the measurements. Force the CPU with -ngl 0 even if a GPU is present.

Terminal — measure
# Télécharger le GGUF une fois (llama-cli -hf le met en cache) puis :
llama-bench -m ~/.cache/llama.cpp/LiquidAI_LFM2-1.2B-GGUF_LFM2-1.2B-Q4_K_M.gguf \
  -ngl 0 -t 8 -p 512 -n 128

# Sortie : une ligne pp512 (prefill) et une ligne tg128 (génération), en tokens/s

# Comparer avec un transformer de taille voisine
llama-bench -m ~/.cache/llama.cpp/Qwen_Qwen3-1.7B-GGUF_Qwen3-1.7B-Q4_K_M.gguf -ngl 0 -t 8 -p 512 -n 128

The orders of magnitude below correspond to what we observe with llama.cpp in Q4_K_M, using all physical cores, on typical machines. They do not replace your own measurements: memory bandwidth (DDR4 versus DDR5, number of channels) can make the result vary twofold between two PCs of the same generation.

Recent laptop (8 cores, DDR5)
LFM2-1.2B: 40 to 70 tokens/s during generation, and several hundred tokens/s during prefill. The 350M exceeds 100 tokens/s. The 2.6B runs at around 25 to 40 tokens/s.
Desktop PC with 4 to 6 cores, DDR4
LFM2-1.2B: 20 to 35 tokens/s, well above reading speed. A Qwen3-1.7B on the same machine is closer to 12 to 20 tokens/s.
Mac Apple Silicon (CPU only)
Comparable to a DDR5 laptop thanks to unified memory; in practice, you’ll let Metal accelerate it, but LFM2 remains comfortable even on CPU alone on a MacBook Air.
Raspberry Pi 5 (8 GB)
LFM2-1.2B: 8 to 12 tokens/s; LFM2-350M: 25 to 35 tokens/s. This is where the gap with a classic transformer is most visible.
LFM2-8B-A1B
With 16 GB of DDR5 RAM, expect 30 to 50 tokens/s: only 1.5B parameters work on each token, but the 5 GB of weights must fit in memory.

The factor that makes the difference in practice is prefilling. On a 1.7B transformer, ingesting a 4,000-token document on the CPU often takes 15 to 30 seconds before the first response; LFM2-1.2B cuts that time by nearly half in tests published by Liquid AI, making local RAG on the CPU genuinely usable rather than merely possible.

!
The thread count is not “all cores”
On a hybrid CPU (Intel Core with P- and E-cores, Apple Silicon), setting -t to the total number of logical cores often reduces speed. Test with only the number of performance cores (for example, -t 8 on an 8P+16E) and compare: the difference can reach 30%.

#Which use cases favor LFM2

LFM2 is not a replacement for Qwen3-8B or Gemma 3 12B. On a GPU with VRAM, a larger transformer will be smarter, period. LFM2 becomes useful as soon as hardware is the constraint: no GPU, limited RAM, a battery to conserve, or a query volume to handle at a fixed cost.

Laptop assistant without a GPU
A smooth chat on a fanless ultraportable, without the fan spinning up. The 1.2B or 2.6B respond faster than you can read.
Batch processing on a CPU server
Classify tickets, extract fields, rewrite product descriptions: thousands of short queries per hour on a VM without a GPU, using the 350M or 700M.
Lightweight local RAG
Fast prefill + 32k context: index notes or internal documentation and answer on CPU in a few seconds.
Embedded and home automation
Raspberry Pi, industrial mini-PC, home automation hub: interpret a transcribed voice command, generate a short response, call a tool.
Mobile
Liquid AI offers its LEAP SDK (Liquid Edge AI Platform) for iOS and Android, along with the Apollo app for testing models on a phone. GGUF files also work in apps based on llama.cpp.
Tool calls
LFM2 instruct variants natively support function calling with dedicated tags, making them an affordable small agent router.

Conversely, stick with a conventional transformer when you have a GPU and VRAM to fill, when the task requires long-form reasoning (math, complex code), or when you depend on a well-stocked fine-tuning ecosystem: Qwen and Llama still have the edge in this area.

i
And compared with SmolLM3, Gemma 3 1B or Qwen3-0.6B?
At the same size, LFM2 is generally a little better on benchmarks and significantly faster on CPU. The tradeoff is the LFM Open License, which is less permissive than Apache 2.0 for a large enterprise, and a younger ecosystem: fewer community fine-tunes and fewer real-world reports.

#Limitations and troubleshooting

“unknown model architecture: lfm2” error
Your engine is too old. Update Ollama, LM Studio, or recompile llama.cpp from a version released after July 2025.
Responses that get stuck in a loop
Lower the temperature to 0.3 and enable min_p 0.15 and repeat_penalty 1.05, the values recommended by Liquid AI. Small models are sensitive to the default settings.
Disappointing speed
Check with ollama ps that the model is fully loaded, reduce the thread count to the performance cores, and close applications that saturate memory bandwidth (a browser with dozens of tabs, for example).
Bad French
The 350M remains limited in French; move up to the 1.2B or 2.6B, trained with a more useful multilingual component.
Overly aggressive quantization
With a 350M or 700M model, Q4 damages quality more than it does on a 7B model. Prefer Q8_0 for smaller sizes: the file remains tiny (under 800 MB for the 700M).
Enterprise licensing
Beyond the revenue threshold set by the LFM Open License, commercial use requires an agreement with Liquid AI. Check before putting it into production.

#Go further

LFM2 makes the most sense in a setup without a GPU or on a small machine. These site guides complement this one:

Local LLM without a GPU (CPU)
Recommended models by RAM capacity and CPU tokens/s benchmarks, to position LFM2 against the alternatives.
Import a GGUF model from Hugging Face into Ollama
Everything about hf.co syntax, Modelfiles, and aliases, useful for locking in the recommended LFM2 parameters.
LLM on Raspberry Pi 5
The quintessential embedded use case, where LFM2’s speed makes the difference.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.