IBM’s Granite 4 locally: installation and use cases d'usage
IBM is quietly developing one of the most interesting open-weight LLM families for businesses: Granite. This guide shows how to install Granite 4 locally with Ollama, explains the hybrid Mamba architecture that reduces memory requirements, details the Apache 2.0 license with no hidden catch, and targets use cases—RAG, function calling, and business tasks—where Granite truly stands out from consumer models.
#Why Granite 4 instead of another model
Granite is not designed to top chatbot rankings. IBM is targeting a different goal: reliable, memory-efficient, legally clean models built to be integrated into enterprise applications. You don't choose Granite to discuss philosophy, but to connect a model to a document repository, have it call tools, or run it at lower cost on modest hardware.
Three things set Granite 4 apart from the rest of the open-weight field. First, a hybrid Mamba architecture that significantly reduces memory usage, especially with long contexts. Second, a strict Apache 2.0 license, without the usage restrictions imposed by Llama or Gemma. Finally, a focus on traceability and governance—signed models, documented training data, and ISO 42001 certification—that IT departments and legal teams appreciate. It is a model designed to pass a compliance review, not just a benchmark.
#The hybrid Mamba architecture and its memory savings
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The real technical argument for Granite 4 is its architecture. A classic transformer relies entirely on attention: with each new token, the model rereads the entire context, and it stores a “KV cache” that grows linearly with the conversation length. The longer the context, the more VRAM usage explodes and the more inference slows down. This is the well-known limitation of large local contexts.
Mamba belongs to another family: state space models. Instead of rereading everything, a Mamba layer maintains a fixed-size recurrent state that summarizes the past. As a result, memory doesn't grow with context length, and throughput remains stable even on very long documents. Mamba's sole drawback is less precise recall of distant details.
Granite 4 combines both. Most blocks are Mamba-2 layers (efficient), interspersed with a minority of attention blocks (precise)—a ratio of about nine Mamba layers to one attention layer. This preserves most of the memory savings while retaining attention’s ability to retrieve an exact detail. In practice, Granite 4 uses significantly less RAM/VRAM than a transformer of comparable size, especially once the context exceeds a few thousand tokens.
#Micro, Tiny, and Small: which variant should you choose
Granite 4 comes in several sizes, with suffixes indicating the target hardware. Versions marked “H” use the hybrid architecture; a non-hybrid version is provided for environments that do not yet support Mamba.
- Granite 4 Micro (~3B, dense)
- The smallest, dense, hybrid model. Ideal for CPU, limited VRAM, or edge use. Fits in ~2 GB in Q4. Perfect for extraction, classification, and lightweight RAG.
- Granite 4 Tiny-H (~7B, MoE)
- Hybrid mixture of experts: many total parameters but few active at each token, making it fast. A good quality/speed compromise on an entry-level card.
- Granite 4 Small-H (~32B, MoE)
- Accessible high end: near-large-model quality for business tasks while retaining the hybrid model's reduced memory footprint. Targeted at workstations or servers.
- Non-hybrid variant
- A classic transformer version exists for runtimes that do not yet support Mamba. Reserve it for compatibility cases—you lose the memory savings.
To get started, Micro is enough to validate use cases and runs everywhere. Move to Tiny-H or Small-H once you know what you want to put into production and have measured a quality shortfall.
#Requirements and VRAM
The software prerequisite boils down to having Ollama installed and up to date—the daemon listens on http://localhost:11434 by default. Make sure you have a recent version: support for Granite 4's Mamba layers requires a fairly recent Ollama, otherwise the model will refuse to load.
- Micro (3B) in Q4
- ≈ 2 GB of VRAM. It also runs on CPU with 8 GB of RAM, slowly but reliably. A RTX 3060 12GB is very comfortable.
- Tiny-H (7B MoE) in Q4
- ≈ 5 GB of VRAM for the weights, but partial MoE activation makes inference lightweight. An 8–12 GB card is more than sufficient.
- Small-H (32B MoE) in Q4
- ≈ 19 GB of VRAM. RTX 4090 24GB or a Mac with 32–48 GB of unified memory. The hybrid approach limits memory blowup with long context.
- Headroom for context
- With Mamba, the KV cache is much lighter than that of an equivalent transformer: you can target large contexts without doubling your VRAM.
#Install Granite 4 via Ollama
Installation follows the usual Ollama workflow. The model is published in the official library under the name granite4; variants are selected with tags. Check the exact tag names on ollama.com/library before pulling a specific size, because they change from one release to the next.
- 01Check OllamaConfirm that the daemon is running and up to date. An older version doesn’t know about Granite 4’s Mamba layers.
- 02Pull the modelDownload the selected variant with ollama pull. Start with Micro to test quickly without saturating your disk or GPU.
- 03Start a first chatollama run ouvre une session interactive. Posez une question métier — résumé d'un texte, extraction de champs — plutôt qu'une devinette généraliste.
- 04Connect an interfacePoint Open WebUI or LM Studio at the local endpoint for a comfortable experience, or call the HTTP API directly from your application.
For application use, Ollama exposes an OpenAI-compatible API on the same port. You can therefore reuse any existing client by simply changing the base URL and model name.
#Apache 2.0, the no-surprises license for business
This is something legal teams review before the benchmarks. Granite 4 is released under the Apache 2.0 license, a permissive, proven open-source license. You can use it commercially, modify it, redistribute it, and integrate it into a closed product, with no user threshold or acceptable-use clause to monitor.
The distinction matters compared with popular alternatives. Meta's “Llama Community License” is not a true open-source license: it imposes restrictions beyond 700 million monthly users and prohibits certain uses. The terms of Gemma (Google) also govern use through an acceptable-use policy. Apache 2.0 has none of that: it is a standard license that legal teams already know and approve without negotiation.
- Commercial use
- Allowed without restrictions on size or sector. No user threshold to monitor.
- Modification and fine-tuning
- You can adapt the model and keep your weights private, with no obligation to publish them.
- Redistribution
- Can be integrated into a proprietary product while retaining the license notice.
- IBM governance
- Cryptographically signed models, documented data, and ISO 42001 certification—concrete arguments for a compliance review.
#RAG and function calling: Granite's strengths
This is where Granite justifies its existence. IBM specifically trained these models for two enterprise use cases: RAG (answering from provided documents) and function calling (reliably calling tools). These are exactly the building blocks needed to build a useful business assistant, not just a chatbot.
For RAG, Granite closely follows the supplied context and limits hallucinations: it tends to rely on the injected passages rather than embellishing. Combined with the hybrid architecture, which handles long documents without exhausting VRAM, this makes it a good engine for querying a local document database. The Micro version is often sufficient, making RAG accessible on modest hardware.
When it comes to tools, Granite produces well-formed function calls and follows the expected JSON schemas. This is the foundation of an agent: the model decides to call a function, reads the result, and continues. Combined with an affordable long context, this delivers reliable business automations without depending on the cloud.
- Document RAG
- Clear strength: good context retention, a low hallucination rate, and a long context that is memory-efficient. Ideal for an internal knowledge base.
- Function calling
- Reliable tool calls and valid JSON — the building block for local agents and integrations with an existing system.
- Structured extraction
- Turn free-form text into clean fields (dates, amounts, entities). The Micro version already does this very well.
- General-purpose conversation
- Respectable without being exceptional. That is not where Granite aims to win.
#Troubleshooting
- The model refuses to load (architecture error)
- Your Ollama is too old for Granite 4's Mamba layers. Update Ollama to a recent version, then run the pull again.
- “model not found” on pull
- The tag does not exist under this name. Check the exact spelling on ollama.com/library—the variant tags change from one release to the next.
- Off-topic RAG responses
- The context is poorly formatted or too noisy. Structure the injected passages, reduce their number, and explicitly ask for answers based only on the documents provided.
- Malformed tool calls
- The JSON schema is ambiguous. Simplify the definition, make required fields explicit, and test first with a single tool before stacking several.
- Speed collapses on Small-H
- The model spills into RAM because it lacks enough VRAM. “ollama ps” shows the GPU/CPU split. Switch to Tiny-H or Micro, or drop down one quantization level.
- “Connection refused”
- The Ollama daemon isn't running. Check with “ollama ps” that it's listening on http://localhost:11434.
#Go further
Granite builds on components already covered on the site. These guides build on this one:
- Install Ollama on Linux
- The prerequisite if the daemon is not yet in place: an install script, a systemd service, and NVIDIA/AMD GPU configuration.
- Choose your quantization (Q4, Q5, Q8, FP16)
- For balancing quality and VRAM on Granite Small-H, where each quantization level changes what fits on your card.
- Local AI in the enterprise: GDPR, sovereignty, and deployment
- The natural complement to the Apache 2.0 license: how to deploy Granite compliantly in an organization.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.