Beginner 10 minMobile

Local LLMs on Android: PocketPal, MLC Chat (2026)

A local LLM on Android is an assistant that responds offline, without the cloud, directly on your phone's processor. Quantized models with 1 to 4 billion parameters now fit in the RAM of a midrange smartphone and respond at a usable speed. This guide covers three approaches—PocketPal and MLC Chat to get started without the command line, and llama.cpp through Termux to go further—with the real constraints of speed, heat, and battery life.

By Léa B.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run a local LLM on Android

Running a local LLM on Android addresses three concrete needs: complete privacy (your prompts never leave the device), offline operation (planes, dead zones, expensive roaming), and free usage (no subscription or per-token billing). The tradeoff is size: a phone runs much smaller models than a PC, so they’re less capable. But for summarization, rewriting, translation, or a simple conversation, a 3B model is more than sufficient.

Privacy
The model runs on the phone's SoC. No data is sent over the Internet, which you can verify by turning off Wi-Fi and mobile data.
Offline
Once the model is downloaded, everything works in airplane mode. Convenient when traveling or in areas without network coverage.
Free to use
No account, no quota, no per-token cost. You only pay for the battery power consumed.
A limitation to know
A smartphone realistically tops out at around 4B parameters. For complex reasoning or code, a local LLM on a PC is still far superior.
i
Local on a phone ≠ local on a PC
Don't compare a 3B model on mobile with GPT-4. Compare it with what an offline assistant can do: rephrase a message, summarize pasted text, or answer a general-knowledge question. In that context, it's impressive. Beyond that, its limitations show quickly.

#What kind of phone you need

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The determining factor is RAM, exactly as on a PC, where VRAM caps the model size. A 3B Q4-quantized model uses about 2 GB, plus the system and application memory. In practice, you need headroom because Android aggressively closes applications that consume too much.

6 GB RAM
The viable minimum. 1B to 3B models in Q4. Close other applications before starting an inference.
8 GB RAM
Comfortable for 3B, possible with 4B. The sweet spot in the latest mid-range segment.
12 GB RAM and up
Recent flagships. Easily handles 4B, and even aggressively quantized 7B, albeit slowly.
Storage
Allow 2 to 5 GB of free space per downloaded model. GGUF files are large.
SoC
A recent Snapdragon 8-series or equivalent delivers a noticeable speed boost. Entry-level chips work, but remain slow.
→
Check your actual RAM
The advertised RAM is not all available: the system reserves part of it. On an 8 GB phone, expect 5 to 6 GB to be realistically allocable to an application. That is the usable margin for loading a model.

#PocketPal: a local LLM in 5 minutes

PocketPal AI is the simplest app to get started. It's an open-source project available on the Play Store that bundles llama.cpp and a catalog of models downloadable directly from Hugging Face, with no command line required.

  1. 01
    Install the application
    Search for « PocketPal AI » in the Google Play Store and install it. The app is free and open source (code available on GitHub).
  2. 02
    Open the model catalog
    In the Models tab, browse the list of recommended models. PocketPal shows the file size and flags those suited to your RAM.
  3. 03
    Download a 1-4B model
    Choose a small model to start with, such as a Qwen 3.5 2B (~1.9 GB) or a Granite 4.2 3B (~2.2 GB) with Q4 quantization. The download occurs over the network, only once.
  4. 04
    Load and chat
    Click Load next to the downloaded model, wait for it to load into memory, then open the chat. The first response starts after a few seconds of warm-up.

PocketPal also lets you adjust inference parameters (temperature, context size, number of CPU threads) and import your own GGUF files if you want a model that isn’t in the catalog. For a first experience with a local LLM on Android, it’s the shortest path.

→
Disable the network to test
Once the model is loaded, turn on airplane mode and start a conversation. It works: concrete proof that inference is 100% local, with no network calls.

#MLC Chat and GPU acceleration

MLC Chat is the demo application for the MLC LLM project. It compiles models to take advantage of the mobile GPU through the Vulkan/OpenCL API instead of running everything on the CPU. On a recent SoC, this can improve speed and slightly reduce power consumption compared with pure CPU execution.

  1. 01
    Retrieve the APK
    MLC Chat is not always available on the Play Store. Download the official APK from the MLC LLM project website (llm.mlc.ai) and allow installation from an external source.
  2. 02
    Download a compiled model
    The app offers models already converted to MLC format (Gemma, Qwen, Granite in smaller sizes). Choose one suited to your RAM.
  3. 03
    Start the chat
    Select the model, wait for initialization, then start chatting. MLC displays the speed in tokens/second at the bottom of the screen.
i
Mobile GPU: the performance gain isn't guaranteed
GPU acceleration depends heavily on the SoC and drivers. On some phones, MLC is faster than llama.cpp on the CPU; on others, the opposite is true because of heat and throttling. Test both on your device before drawing conclusions.

#llama.cpp via Termux to go further

For total control—choosing the exact model and quantization and exposing a local server—the advanced path goes through Termux, an Android terminal emulator that provides access to a Linux environment without root. You compile llama.cpp there and run inference from the command line, just as on a Linux PC.

!
Install Termux from F-Droid, not the Play Store
The Play Store version is obsolete and abandoned. Install Termux from F-Droid or GitHub to get up-to-date packages. This is the most common source of errors for beginners.
  1. 01
    Install and prepare Termux
    Install Termux from F-Droid, open it, then update the core packages.
  2. 02
    Install the build tools
    Get the compiler, git, and cmake needed to build llama.cpp.
  3. 03
    Build llama.cpp
    Clone the official repository and compile it. On mobile, the CPU build (ARM NEON) is the most reliable.
  4. 04
    Download a GGUF model
    Download a small .gguf file (1-3B in Q4) from Hugging Face with curl or wget.
  5. 05
    Run inference
    Use llama-cli to chat, or llama-server to expose an OpenAI-compatible API on localhost that is accessible from a browser.
Termux — setup
# Mettre à jour les paquets
pkg update && pkg upgrade -y

# Outils de compilation
pkg install -y git cmake clang wget
Termux—compile llama.cpp
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j$(nproc)
Termux — download a model and chat
# Exemple : un petit modèle Q4 depuis Hugging Face
wget -O qwen3.5-2b-q4.gguf "<URL_DU_FICHIER_GGUF>"

# Discuter en ligne de commande
./build/bin/llama-cli -m qwen3.5-2b-q4.gguf -p "Résume ce texte : ..." -n 256
Termux — local OpenAI-compatible server
# Expose une API sur http://localhost:8080
./build/bin/llama-server -m qwen3.5-2b-q4.gguf --port 8080

The server mode is useful: it exposes an OpenAI-compatible endpoint on localhost that you can query from your phone's browser or another app. It's the same approach as the Ollama server on a PC (which listens on port 11434 by default), transferred to your pocket.

#Which 1–4B models actually run

Everything comes down to the 1- to 4-billion-parameter range, in Q4_K_M quantization (the best quality/size compromise, as on PC). Above 4B, a phone slows to a crawl or runs out of RAM. Here are the families that deliver the best results in French.

Qwen 3.5 2B
The default small model for 2026 (~1.9 GB in Q4). Excellent for multilingual and French use, with extended context and an Apache 2.0 license. The best starting point for speed on a 6 GB phone.
Qwen 3.5 4B
The quality boost (~3.4 GB in Q4) without leaving the mobile tier. Consistent and capable in Q4 if your phone has 8 GB+. Apache 2.0.
Granite 4.2 3B
IBM’s small model (~2.2 GB in Q4), highly memory-efficient and token-efficient. A solid offline general-purpose model when RAM is limited. Apache 2.0.
Gemma 4 E2B
Google's compact variant (~4.3 GB), multimodal and back under the Apache 2.0 license since April 2026. Best reserved for phones with 8 GB or more.
→
Memory footprint in Q4
Rule of thumb: a 1B model fits in ~1 GB, a 3B model in ~2 GB, and a 4B model in ~3 GB in Q4. Add system and app memory: that's why an 8 GB phone targets 3–4B and not beyond.

#Speed, heat, and battery life: what you need to know

LLM inference fully loads the CPU (or GPU), which has three very concrete consequences on a phone: speed remains modest, the device heats up, and the battery drains quickly during generation. Nothing prohibitive, but you need to keep the general scale in mind.

Speed
Roughly expect 5 to 15 tokens/second for a 3B Q4 on a mid-to-high-end SoC. It is readable in real time, but nowhere near a PC GPU. A 1B is considerably faster.
Heat and throttling
After a long generation, the SoC heats up and throttles, slowing subsequent responses. Let it cool down between long requests.
Battery
A sustained session can drain several percent of the battery in a few minutes. For occasional use, it’s negligible; for intensive use, plug in the charger.
Context
The larger the context window, the more memory and processing time increase. On mobile, keep the context reasonable (2k-4k tokens) to stay responsive.
!
Do not charge using fast charging
Running an LLM while fast-charging combines two sources of heat. Over time, repeated heating wears down the battery. For intensive use, prefer slow charging or take breaks.

#Troubleshooting

The app closes while loading the model
Not enough RAM. Close all other applications, choose a smaller model (1B instead of 3B), or use a lighter quantization.
Very slow responses
Reduce the context size, lower the number of generated tokens, or switch to a smaller model. Also check that the phone is not throttling because of heat.
Termux: command not found after compilation
The build failed silently. Re-run the compilation and read the errors. Make sure git, cmake, and clang are installed.
Termux abandoned / outdated packages
You installed the Play Store version. Uninstall it and reinstall Termux from F-Droid or GitHub.
Inconsistent responses
Normal for a very small model on a complex task. Simplify the request, or accept that a 1-3B model has reasoning limitations.

#Go further

Once you are comfortable on mobile, you will probably want a more capable machine at home for demanding tasks. These site guides build on this one:

Run an LLM locally without a GPU (CPU only)
The same logic as on mobile, but on PC: models recommended by RAM capacity and tips for accelerating pure-CPU operation.
Choose your quantization (Q4, Q5, Q8, FP16)
Understand why Q4_K_M is the recommended compromise, on both phones and PCs.
Install Ollama
For a real LLM server at home, accessible from your phone over the local network.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.