Intermediate 10 minFalcon

Falcon H1 locally: the Mamba hybrid from the United Arab Emirates

Falcon H1 is the Technology Innovation Institute of Abu Dhabi’s open-weight model family, and the first TII family to combine standard attention and Mamba layers in every block. This guide explains what this hybrid architecture changes, which Falcon H1 size to install locally based on your VRAM, how to measure its memory savings on long contexts, and whether it’s worth replacing Mistral Small or Qwen3 in your Ollama stack.

By Samir K.·Update 2026-09-27·Tested on Windows, macOS, and Linux

#Why take an interest in Falcon H1

The first Falcon models, in 2023, made their mark among open-weight models before being overtaken by Llama, Mistral, and then Qwen. Falcon H1, released in May 2025 by Abu Dhabi’s Technology Innovation Institute (TII), takes a different approach: instead of chasing conventional transformers, TII starts with a different architecture that combines attention and state-space models (Mamba-2) to produce compact models that hold their own against competitors twice their size.

For local use, its appeal comes down to three points. The lineup covers every configuration, from 0.5 billion parameters for a Raspberry Pi or an old laptop up to 34 billion for a RTX 4090 or a Mac with unified memory. The advertised context reaches 256,000 tokens, and the hybrid architecture makes it genuinely usable locally where a conventional transformer saturates VRAM. Finally, Falcon H1 was natively trained on 18 languages, including French and Arabic, making it a serious candidate for French-language text without switching to a model focused on English or Chinese.

i
Falcon H1 and Falcon H1R
This guide covers the base Falcon H1 family (Instruct). Falcon H1R 7B, released in early 2026, is the step-by-step reasoning variant built on the same architecture. Everything related to installation and memory applies to both; only response behavior differs, with H1R being slower because it is more verbose.

#The attention-Mamba hybrid in two words

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A classic transformer relies entirely on attention. With each new token, the model rereads the entire context and retains a pair of vectors for every previous token—the famous KV cache. This memory grows linearly with the length of the conversation or document. It, not the model weights, is what overflows your graphics card when you load a long report.

Mamba belongs to another family: state space models. A Mamba layer doesn't reread the past; it maintains a fixed-size recurrent state that summarizes what came before, like a modernized recurrent network. Memory per token is constant, and throughput remains stable regardless of context length. The known trade-off is less precise recall of distant details: a pure Mamba model is less able to retrieve a specific number buried on page 40.

Hybrids aim for the best of both worlds. Where IBM Granite 4 or Jamba alternate Mamba layers and attention layers, Falcon H1 uses a parallel design: in each block, attention heads and Mamba-2 heads work side by side on the same input, and their outputs are concatenated before the next layer. TII tuned the ratio in favor of the Mamba channels, keeping attention memory far below that of an equivalently sized transformer while retaining enough attention to retrieve an exact detail.

Classic transformer (Mistral, Qwen, Llama)
100% attention. Maximum retrieval quality, but KV cache is proportional to the context: VRAM soars beyond a few tens of thousands of tokens.
Pure Mamba (Falcon Mamba 7B)
Constant memory per token, very fast on long streams. Weaker recall of distant details, and software support is still uneven.
Sequential hybrid (Granite 4, Jamba)
Mostly Mamba layers, interspersed with attention layers. Good memory savings and a simple architecture to implement in runtimes.
Parallel hybrid (Falcon H1)
Attention and Mamba-2 in every block, with concatenated outputs. At each layer, the model decides what to entrust to precise memory or compressed memory.
→
What this changes in practice
With a 2,000-token prompt, you won’t see any difference from a standard transformer. Load a 60-page contract or a three-hour conversation, and the gap becomes visible: Falcon H1 maintains an almost flat memory footprint, whereas Mistral Small or Qwen3 must quantize their KV cache or reduce the context.

#Available sizes: from 0.5B to 34B

TII releases Falcon H1 in six sizes, each available in Base (pretrained) and Instruct (chat) versions. Only the Instruct versions matter to you for Ollama or LM Studio use. They all share the same architecture and multilingual tokenizer.

Falcon H1 0.5B
The smallest one. For embedded use, pipeline testing, or a Raspberry Pi. Don't expect it to handle serious writing tasks.
Falcon H1 1.5B
Classification, simple extraction, autocomplete. Runs on any recent CPU with less than 2 GB of memory.
Falcon H1 1.5B-Deep
The same number of parameters as the 1.5B, but many more, narrower layers. TII positions it alongside models with 7 to 10 billion parameters. This is the variant to try on a laptop without a GPU.
Falcon H1 3B
The trade-off for a 4–6 GB card or a 16 GB Mac. Summarization, French chat, lightweight RAG.
Falcon H1 7B
The mainstream option. Competes with Qwen3 8B and outperforms the previous generation's 7B models. A RTX 3060 12GB runs it with a generous context.
Falcon H1 34B
The high-end option. TII compares it with Qwen3 32B, Gemma 3 27B, and Llama 4 Scout. It requires at least 24 GB of VRAM in Q4, or a 48 GB unified-memory Mac for comfortable use.

The 1.5B-Deep variant deserves a mention. TII bet on a narrow but very deep model, with nearly three times as many layers as the standard 1.5B. The result is slower per token because each layer runs sequentially, but significantly more capable. On CPU, where memory bandwidth limits everything, it is often the best quality-per-gigabyte ratio in the entire lineup.

#Requirements and VRAM

On the software side, you need a recent Ollama. Support for the falcon-h1 architecture was added to llama.cpp in summer 2025, and Ollama inherited it in subsequent releases. An earlier Ollama will refuse to load the model with an unknown architecture error. The daemon listens by default on http://localhost:11434; Open WebUI or LM Studio connect to it without any special configuration.

0.5B and 1.5B in Q4_K_M
Less than 1.5 GB. CPU only with 4 GB of free RAM. No GPU required.
3B in Q4_K_M
≈ 2 GB of VRAM. Any card with 4 GB or more, or 8 GB of RAM in CPU mode.
7B in Q4_K_M
≈ 5 GB of VRAM for the weights. Comfortable with RTX 3060 12GB or RTX 4070 12GB, with room for a 32K-token context or more.
7B in Q8_0
≈ 8 GB. For a RTX 4080 16GB or a 24 GB Mac if you want maximum accuracy for extraction.
34B in Q4_K_M
≈ 20 GB. RTX 4090 is just enough at 24 GB; keep an eye on context. M4 Pro 48 GB or Mac Studio will be much more comfortable.
34B in Q8_0
≈ 36 GB. Reserved for Macs with 64 GB or more, or a dual-GPU setup.
i
Quantization reference
Q4_K_M remains the recommended default, offering the best quality-to-size ratio. Move to Q5_K_M or Q8_0 only if VRAM allows it and you measure a difference on your tasks. On Falcon H1, the headroom freed by the reduced KV cache is often worth more than one additional quantization level.

#Install Falcon H1 locally based on its VRAM

TII publishes official GGUFs for every size on Hugging Face, in repositories named tiiuae/Falcon-H1-<taille>-Instruct-GGUF. Ollama can pull a GGUF directly from Hugging Face using the hf.co prefix, with the desired quantization specified after the colon. While you’re there, check ollama.com/library to see whether an official falcon-h1 tag has appeared since this guide was written; if so, prefer it because it includes a validated chat template.

  1. 01
    Update Ollama
    Rerun the official installer or your package manager, then check the version. This is the step everyone skips, and it explains most loading failures.
  2. 02
    Choose the size based on VRAM
    12 GB or less: 7B in Q4_K_M. 4 to 6 GB: 3B. No GPU: 1.5B-Deep. 24 GB or Mac 48 GB: 34B. Don’t pull several sizes at once; each GGUF weighs 1 to 20 GB.
  3. 03
    Pull the GGUF from Hugging Face
    Use ollama pull with the repository’s hf.co path and quantization tag. Ollama downloads the file, reads its metadata, and generates the chat template from it.
  4. 04
    Start a session and test in French
    ollama run ouvre un chat interactif. Posez une vraie tâche, résumé d'un texte ou extraction de champs, plutôt qu'une devinette. Vérifiez que le modèle répond en français sans glisser vers l'anglais.
  5. 05
    Control GPU/CPU allocation
    ollama ps indique le pourcentage du modèle chargé sur le GPU. Si une partie est en CPU, réduisez la taille ou la quantification, sinon le débit s'effondre.
Terminal — Falcon H1 7B on a 12 GB card
# 1. Vérifier la version d'Ollama (doit dater d'après l'été 2025)
ollama --version

# 2. Tirer le GGUF officiel TII en Q4_K_M (~4,5 Go)
ollama pull hf.co/tiiuae/Falcon-H1-7B-Instruct-GGUF:Q4_K_M

# 3. Lancer une session interactive
ollama run hf.co/tiiuae/Falcon-H1-7B-Instruct-GGUF:Q4_K_M

# 4. Vérifier que tout est sur le GPU
ollama ps
Terminal — other sizes
# Sans GPU : la variante 1.5B-Deep, étroite mais très profonde
ollama pull hf.co/tiiuae/Falcon-H1-1.5B-Deep-Instruct-GGUF:Q4_K_M

# Carte 4-6 Go ou Mac 16 Go
ollama pull hf.co/tiiuae/Falcon-H1-3B-Instruct-GGUF:Q4_K_M

# RTX 4090 24 Go ou Mac 48 Go
ollama pull hf.co/tiiuae/Falcon-H1-34B-Instruct-GGUF:Q4_K_M

The full name with the hf.co prefix is cumbersome to type and hard to read in Open WebUI. Create a local alias with a Modelfile: you can use it to set the context and a French system prompt, and the model appears under a short name in all your interfaces.

Terminal — short alias with Modelfile
cat > Modelfile.falcon-h1 <<'EOF'
FROM hf.co/tiiuae/Falcon-H1-7B-Instruct-GGUF:Q4_K_M
PARAMETER num_ctx 32768
PARAMETER temperature 0.3
SYSTEM "Tu es un assistant précis. Tu réponds en français, de façon concise."
EOF

ollama create falcon-h1:7b -f Modelfile.falcon-h1
ollama run falcon-h1:7b

For application use, Ollama exposes its HTTP API on the same port, with an OpenAI-compatible endpoint. Any existing client works by changing the base URL and model name.

Terminal — local API call
curl http://localhost:11434/api/chat -d '{
  "model": "falcon-h1:7b",
  "messages": [
    { "role": "user", "content": "Résume ce texte en trois points : ..." }
  ],
  "options": { "num_ctx": 32768 },
  "stream": false
}'
!
llama.cpp and LM Studio users
llama.cpp loads the same GGUF files with the -hf option, for example llama-server -hf tiiuae/Falcon-H1-7B-Instruct-GGUF:Q4_K_M. LM Studio uses the same engine: update it before looking for Falcon H1 in its catalog, because an older version will not list the file as compatible.

#The measured memory advantage on long contexts

This is Falcon H1’s central promise when run locally, and it can be verified in ten minutes. The protocol is simple: load the same model with two context sizes and compare the memory usage reported by ollama ps. With a conventional transformer, going from 8K to 64K tokens increases memory usage by several gigabytes. With Falcon H1, the increase is still there because the attention component retains a KV cache, but it is significantly smaller.

Terminal — measuring context cost
# Contexte 8K : noter la colonne SIZE de ollama ps
OLLAMA_CONTEXT_LENGTH=8192 ollama run falcon-h1:7b "Bonjour" && ollama ps

# Décharger, puis recharger avec 64K
ollama stop falcon-h1:7b
OLLAMA_CONTEXT_LENGTH=65536 ollama run falcon-h1:7b "Bonjour" && ollama ps

# Même exercice avec un transformeur classique pour comparer
ollama stop falcon-h1:7b
OLLAMA_CONTEXT_LENGTH=65536 ollama run qwen3:8b "Bonjour" && ollama ps

Two clarifications for reading the figures. First, Ollama reserves the KV cache in advance for the entire requested context: the displayed memory therefore reflects the reserved capacity, not what your prompt actually uses. Second, if you enabled KV-cache quantization in the Ollama configuration, it also applies to Falcon H1’s attention component, further reducing the measured gap but not changing the conclusion: with the same VRAM, Falcon H1 accepts a much longer context than a transformer of the same size.

In practice, on a 12 GB card, Falcon H1 7B in Q4_K_M leaves room for a 64K-token context without spilling onto the CPU, whereas Qwen3 8B or Mistral 7B require quantizing the KV cache or limiting the context to around 32K. Generation throughput degrades much less as the context fills: the Mamba component processes each new token in constant time.

!
Long context does not mean perfect recall
Falcon H1 accepts 256K tokens, but accepting is not understanding. Like any model, its accuracy drops on details buried deep in the prompt, and the Mamba component contributes to this. To retrieve an exact figure from 200 pages, RAG with chunking and search remains more reliable than a giant context. Falcon H1’s long context shines for summarization, synthesis, and conversation tracking—not needle-in-a-haystack searches.

#Mistral Small vs. Qwen3: the verdict

The question everyone is asking: should you replace your usual model? The answer depends on the size, because Falcon H1 is in a different category depending on the variant.

Falcon H1 7B versus Qwen3 8B
Overall quality is comparable. Qwen3 retains an edge in coding and structured reasoning, while Falcon H1 is more natural in French and especially much more efficient as the context grows. For document summarization and French-language chat on 12 GB, Falcon H1 comes out ahead. For coding, stick with Qwen3.
Falcon H1 7B versus Mistral Small 3.2
This is a different class: Mistral Small has 24 billion parameters and requires 14 to 15 GB in Q4. If you have the VRAM, Mistral Small remains better for most tasks. Falcon H1 7B is the choice when the card has 12 GB or the context must exceed 32K.
Falcon H1 34B versus Mistral Small 3.2
The interesting matchup. Falcon H1 34B is stronger on knowledge and reasoning benchmarks and handles very long contexts better, but it costs 5 GB more in Q4 and has no vision. Mistral Small fits on a 16 GB card, reads images, is under Apache 2.0, and benefits from a more mature ecosystem, especially for function calling.
Falcon H1 34B versus Qwen3 32B
At comparable VRAM, Qwen3 32B remains the reference for coding and agents. Falcon H1 34B is preferable for long documents in French or Arabic, and on Mac, where unified memory handles its 20 GB effortlessly.

The verdict in one sentence: Falcon H1 is the best model to install when memory is your constraint on long contexts, or when French and Arabic are your working languages. It is not the best general-purpose all-around model, and its ecosystem is not as mature as Mistral or Qwen. On a workstation with 24 GB, Mistral Small remains the safe choice for a daily assistant; Falcon H1 34B is the one to add alongside it for large documents.

→
The right test before deciding
Take three real documents from your work, 20 to 80 pages long, and ask Falcon H1 7B and your current model the same series of questions, with the same context. Compare the quality of the summaries, the accuracy of the quoted figures, and the throughput shown at the end of the response with ollama run's --verbose option. It's more revealing than any public benchmark.

#Falcon-LLM license: read before deploying

Falcon H1 is released under TII's Falcon-LLM license. It is derived from Apache 2.0 and allows commercial use, modification, and redistribution, but adds an acceptable-use policy and some attribution requirements. It does not offer Apache 2.0's full freedom, as provided by Mistral Small or Granite, nor is it a restrictive license in the sense of Llama with its user thresholds.

For personal or internal use, the question doesn't arise. For a redistributed commercial product, take ten minutes to read the text on the TII website and verify that your use case isn't among the excluded uses. The site's catalog lists the license for each Falcon entry, and it may change from one generation to the next.

#Troubleshooting

Unknown model architecture falcon-h1 error
Your Ollama, LM Studio, or llama.cpp is too old to recognize the hybrid architecture. Update it, restart the daemon, and rerun the pull. There is no other workaround: Mamba-2 layers cannot load without engine support.
The hf.co pull fails or cannot find the tag
Check the exact spelling of the repository and quantization on the Hugging Face page, especially the capitalization of Q4_K_M. If the repository doesn't offer the requested tag, open the Files tab to view the list of available GGUF files.
Responses in English even though you write in French
Add a system prompt that enforces the language, as in the Modelfile above. Falcon H1's multilingual tokenizer handles French very well, but without instructions, the Instruct model sometimes follows the dominant language in its training data.
Very low throughput on 34B
ollama ps montre probablement une répartition partielle sur le CPU. En 24 Go, réduisez le contexte à 16K ou passez en Q4_K_S. Sur Mac, augmentez la limite de mémoire GPU allouable si le système en réserve trop.
Higher-than-expected memory usage with long context
Normal: Ollama reserves the attention KV cache for the entire declared context, even when empty. Always compare with another model at the same context length, and enable KV-cache quantization if you want to save a little more.
Incorrect chat template, visible tags in responses
Ollama generates the template from the GGUF metadata. If tags appear, pull the official TII GGUF instead of a third-party conversion, or define TEMPLATE explicitly in your Modelfile.

#Go further

Falcon H1 makes full sense once you understand the context and memory concepts it uses. These site guides complement this one:

Understanding the context window
What 8K, 32K, or 256K tokens represent, what they cost in VRAM, and why the KV cache is the real bottleneck in conventional transformers.
Quantify the KV cache to save VRAM
The other lever for extending context, which can be combined with Falcon H1's hybrid architecture.
IBM's Granite 4 locally
The other Mamba hybrid of the moment, with a sequential architecture and under Apache 2.0. Comparing it with Falcon H1 helps clarify TII's design choices.
Choose your quantization (Q4, Q5, Q8, FP16)
To choose between Q4_K_M and Q8_0 for each Falcon H1 size based on your VRAM.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.