Falcon H1 locally: the Mamba hybrid from the United Arab Emirates
Falcon H1 is the Technology Innovation Institute of Abu Dhabi’s open-weight model family, and the first TII family to combine standard attention and Mamba layers in every block. This guide explains what this hybrid architecture changes, which Falcon H1 size to install locally based on your VRAM, how to measure its memory savings on long contexts, and whether it’s worth replacing Mistral Small or Qwen3 in your Ollama stack.
#Why take an interest in Falcon H1
The first Falcon models, in 2023, made their mark among open-weight models before being overtaken by Llama, Mistral, and then Qwen. Falcon H1, released in May 2025 by Abu Dhabi’s Technology Innovation Institute (TII), takes a different approach: instead of chasing conventional transformers, TII starts with a different architecture that combines attention and state-space models (Mamba-2) to produce compact models that hold their own against competitors twice their size.
For local use, its appeal comes down to three points. The lineup covers every configuration, from 0.5 billion parameters for a Raspberry Pi or an old laptop up to 34 billion for a RTX 4090 or a Mac with unified memory. The advertised context reaches 256,000 tokens, and the hybrid architecture makes it genuinely usable locally where a conventional transformer saturates VRAM. Finally, Falcon H1 was natively trained on 18 languages, including French and Arabic, making it a serious candidate for French-language text without switching to a model focused on English or Chinese.
#The attention-Mamba hybrid in two words
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A classic transformer relies entirely on attention. With each new token, the model rereads the entire context and retains a pair of vectors for every previous token—the famous KV cache. This memory grows linearly with the length of the conversation or document. It, not the model weights, is what overflows your graphics card when you load a long report.
Mamba belongs to another family: state space models. A Mamba layer doesn't reread the past; it maintains a fixed-size recurrent state that summarizes what came before, like a modernized recurrent network. Memory per token is constant, and throughput remains stable regardless of context length. The known trade-off is less precise recall of distant details: a pure Mamba model is less able to retrieve a specific number buried on page 40.
Hybrids aim for the best of both worlds. Where IBM Granite 4 or Jamba alternate Mamba layers and attention layers, Falcon H1 uses a parallel design: in each block, attention heads and Mamba-2 heads work side by side on the same input, and their outputs are concatenated before the next layer. TII tuned the ratio in favor of the Mamba channels, keeping attention memory far below that of an equivalently sized transformer while retaining enough attention to retrieve an exact detail.
- Classic transformer (Mistral, Qwen, Llama)
- 100% attention. Maximum retrieval quality, but KV cache is proportional to the context: VRAM soars beyond a few tens of thousands of tokens.
- Pure Mamba (Falcon Mamba 7B)
- Constant memory per token, very fast on long streams. Weaker recall of distant details, and software support is still uneven.
- Sequential hybrid (Granite 4, Jamba)
- Mostly Mamba layers, interspersed with attention layers. Good memory savings and a simple architecture to implement in runtimes.
- Parallel hybrid (Falcon H1)
- Attention and Mamba-2 in every block, with concatenated outputs. At each layer, the model decides what to entrust to precise memory or compressed memory.
#Available sizes: from 0.5B to 34B
TII releases Falcon H1 in six sizes, each available in Base (pretrained) and Instruct (chat) versions. Only the Instruct versions matter to you for Ollama or LM Studio use. They all share the same architecture and multilingual tokenizer.
- Falcon H1 0.5B
- The smallest one. For embedded use, pipeline testing, or a Raspberry Pi. Don't expect it to handle serious writing tasks.
- Falcon H1 1.5B
- Classification, simple extraction, autocomplete. Runs on any recent CPU with less than 2 GB of memory.
- Falcon H1 1.5B-Deep
- The same number of parameters as the 1.5B, but many more, narrower layers. TII positions it alongside models with 7 to 10 billion parameters. This is the variant to try on a laptop without a GPU.
- Falcon H1 3B
- The trade-off for a 4–6 GB card or a 16 GB Mac. Summarization, French chat, lightweight RAG.
- Falcon H1 7B
- The mainstream option. Competes with Qwen3 8B and outperforms the previous generation's 7B models. A RTX 3060 12GB runs it with a generous context.
- Falcon H1 34B
- The high-end option. TII compares it with Qwen3 32B, Gemma 3 27B, and Llama 4 Scout. It requires at least 24 GB of VRAM in Q4, or a 48 GB unified-memory Mac for comfortable use.
The 1.5B-Deep variant deserves a mention. TII bet on a narrow but very deep model, with nearly three times as many layers as the standard 1.5B. The result is slower per token because each layer runs sequentially, but significantly more capable. On CPU, where memory bandwidth limits everything, it is often the best quality-per-gigabyte ratio in the entire lineup.
#Requirements and VRAM
On the software side, you need a recent Ollama. Support for the falcon-h1 architecture was added to llama.cpp in summer 2025, and Ollama inherited it in subsequent releases. An earlier Ollama will refuse to load the model with an unknown architecture error. The daemon listens by default on http://localhost:11434; Open WebUI or LM Studio connect to it without any special configuration.
- 0.5B and 1.5B in Q4_K_M
- Less than 1.5 GB. CPU only with 4 GB of free RAM. No GPU required.
- 3B in Q4_K_M
- ≈ 2 GB of VRAM. Any card with 4 GB or more, or 8 GB of RAM in CPU mode.
- 7B in Q4_K_M
- ≈ 5 GB of VRAM for the weights. Comfortable with RTX 3060 12GB or RTX 4070 12GB, with room for a 32K-token context or more.
- 7B in Q8_0
- ≈ 8 GB. For a RTX 4080 16GB or a 24 GB Mac if you want maximum accuracy for extraction.
- 34B in Q4_K_M
- ≈ 20 GB. RTX 4090 is just enough at 24 GB; keep an eye on context. M4 Pro 48 GB or Mac Studio will be much more comfortable.
- 34B in Q8_0
- ≈ 36 GB. Reserved for Macs with 64 GB or more, or a dual-GPU setup.
#Install Falcon H1 locally based on its VRAM
TII publishes official GGUFs for every size on Hugging Face, in repositories named tiiuae/Falcon-H1-<taille>-Instruct-GGUF. Ollama can pull a GGUF directly from Hugging Face using the hf.co prefix, with the desired quantization specified after the colon. While you’re there, check ollama.com/library to see whether an official falcon-h1 tag has appeared since this guide was written; if so, prefer it because it includes a validated chat template.
- 01Update OllamaRerun the official installer or your package manager, then check the version. This is the step everyone skips, and it explains most loading failures.
- 02Choose the size based on VRAM12 GB or less: 7B in Q4_K_M. 4 to 6 GB: 3B. No GPU: 1.5B-Deep. 24 GB or Mac 48 GB: 34B. Don’t pull several sizes at once; each GGUF weighs 1 to 20 GB.
- 03Pull the GGUF from Hugging FaceUse ollama pull with the repository’s hf.co path and quantization tag. Ollama downloads the file, reads its metadata, and generates the chat template from it.
- 04Start a session and test in Frenchollama run ouvre un chat interactif. Posez une vraie tâche, résumé d'un texte ou extraction de champs, plutôt qu'une devinette. Vérifiez que le modèle répond en français sans glisser vers l'anglais.
- 05Control GPU/CPU allocationollama ps indique le pourcentage du modèle chargé sur le GPU. Si une partie est en CPU, réduisez la taille ou la quantification, sinon le débit s'effondre.
The full name with the hf.co prefix is cumbersome to type and hard to read in Open WebUI. Create a local alias with a Modelfile: you can use it to set the context and a French system prompt, and the model appears under a short name in all your interfaces.
For application use, Ollama exposes its HTTP API on the same port, with an OpenAI-compatible endpoint. Any existing client works by changing the base URL and model name.
#The measured memory advantage on long contexts
This is Falcon H1’s central promise when run locally, and it can be verified in ten minutes. The protocol is simple: load the same model with two context sizes and compare the memory usage reported by ollama ps. With a conventional transformer, going from 8K to 64K tokens increases memory usage by several gigabytes. With Falcon H1, the increase is still there because the attention component retains a KV cache, but it is significantly smaller.
Two clarifications for reading the figures. First, Ollama reserves the KV cache in advance for the entire requested context: the displayed memory therefore reflects the reserved capacity, not what your prompt actually uses. Second, if you enabled KV-cache quantization in the Ollama configuration, it also applies to Falcon H1’s attention component, further reducing the measured gap but not changing the conclusion: with the same VRAM, Falcon H1 accepts a much longer context than a transformer of the same size.
In practice, on a 12 GB card, Falcon H1 7B in Q4_K_M leaves room for a 64K-token context without spilling onto the CPU, whereas Qwen3 8B or Mistral 7B require quantizing the KV cache or limiting the context to around 32K. Generation throughput degrades much less as the context fills: the Mamba component processes each new token in constant time.
#Mistral Small vs. Qwen3: the verdict
The question everyone is asking: should you replace your usual model? The answer depends on the size, because Falcon H1 is in a different category depending on the variant.
- Falcon H1 7B versus Qwen3 8B
- Overall quality is comparable. Qwen3 retains an edge in coding and structured reasoning, while Falcon H1 is more natural in French and especially much more efficient as the context grows. For document summarization and French-language chat on 12 GB, Falcon H1 comes out ahead. For coding, stick with Qwen3.
- Falcon H1 7B versus Mistral Small 3.2
- This is a different class: Mistral Small has 24 billion parameters and requires 14 to 15 GB in Q4. If you have the VRAM, Mistral Small remains better for most tasks. Falcon H1 7B is the choice when the card has 12 GB or the context must exceed 32K.
- Falcon H1 34B versus Mistral Small 3.2
- The interesting matchup. Falcon H1 34B is stronger on knowledge and reasoning benchmarks and handles very long contexts better, but it costs 5 GB more in Q4 and has no vision. Mistral Small fits on a 16 GB card, reads images, is under Apache 2.0, and benefits from a more mature ecosystem, especially for function calling.
- Falcon H1 34B versus Qwen3 32B
- At comparable VRAM, Qwen3 32B remains the reference for coding and agents. Falcon H1 34B is preferable for long documents in French or Arabic, and on Mac, where unified memory handles its 20 GB effortlessly.
The verdict in one sentence: Falcon H1 is the best model to install when memory is your constraint on long contexts, or when French and Arabic are your working languages. It is not the best general-purpose all-around model, and its ecosystem is not as mature as Mistral or Qwen. On a workstation with 24 GB, Mistral Small remains the safe choice for a daily assistant; Falcon H1 34B is the one to add alongside it for large documents.
#Falcon-LLM license: read before deploying
Falcon H1 is released under TII's Falcon-LLM license. It is derived from Apache 2.0 and allows commercial use, modification, and redistribution, but adds an acceptable-use policy and some attribution requirements. It does not offer Apache 2.0's full freedom, as provided by Mistral Small or Granite, nor is it a restrictive license in the sense of Llama with its user thresholds.
For personal or internal use, the question doesn't arise. For a redistributed commercial product, take ten minutes to read the text on the TII website and verify that your use case isn't among the excluded uses. The site's catalog lists the license for each Falcon entry, and it may change from one generation to the next.
#Troubleshooting
- Unknown model architecture falcon-h1 error
- Your Ollama, LM Studio, or llama.cpp is too old to recognize the hybrid architecture. Update it, restart the daemon, and rerun the pull. There is no other workaround: Mamba-2 layers cannot load without engine support.
- The hf.co pull fails or cannot find the tag
- Check the exact spelling of the repository and quantization on the Hugging Face page, especially the capitalization of Q4_K_M. If the repository doesn't offer the requested tag, open the Files tab to view the list of available GGUF files.
- Responses in English even though you write in French
- Add a system prompt that enforces the language, as in the Modelfile above. Falcon H1's multilingual tokenizer handles French very well, but without instructions, the Instruct model sometimes follows the dominant language in its training data.
- Very low throughput on 34B
- ollama ps montre probablement une répartition partielle sur le CPU. En 24 Go, réduisez le contexte à 16K ou passez en Q4_K_S. Sur Mac, augmentez la limite de mémoire GPU allouable si le système en réserve trop.
- Higher-than-expected memory usage with long context
- Normal: Ollama reserves the attention KV cache for the entire declared context, even when empty. Always compare with another model at the same context length, and enable KV-cache quantization if you want to save a little more.
- Incorrect chat template, visible tags in responses
- Ollama generates the template from the GGUF metadata. If tags appear, pull the official TII GGUF instead of a third-party conversion, or define TEMPLATE explicitly in your Modelfile.
#Go further
Falcon H1 makes full sense once you understand the context and memory concepts it uses. These site guides complement this one:
- Understanding the context window
- What 8K, 32K, or 256K tokens represent, what they cost in VRAM, and why the KV cache is the real bottleneck in conventional transformers.
- Quantify the KV cache to save VRAM
- The other lever for extending context, which can be combined with Falcon H1's hybrid architecture.
- IBM's Granite 4 locally
- The other Mamba hybrid of the moment, with a sequential architecture and under Apache 2.0. Comparing it with Falcon H1 helps clarify TII's design choices.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To choose between Q4_K_M and Q8_0 for each Falcon H1 size based on your VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.