Local LLMs on Android: PocketPal, MLC Chat (2026)
A local LLM on Android is an assistant that responds offline, without the cloud, directly on your phone's processor. Quantized models with 1 to 4 billion parameters now fit in the RAM of a midrange smartphone and respond at a usable speed. This guide covers three approaches—PocketPal and MLC Chat to get started without the command line, and llama.cpp through Termux to go further—with the real constraints of speed, heat, and battery life.
#Why run a local LLM on Android
Running a local LLM on Android addresses three concrete needs: complete privacy (your prompts never leave the device), offline operation (planes, dead zones, expensive roaming), and free usage (no subscription or per-token billing). The tradeoff is size: a phone runs much smaller models than a PC, so they’re less capable. But for summarization, rewriting, translation, or a simple conversation, a 3B model is more than sufficient.
- Privacy
- The model runs on the phone's SoC. No data is sent over the Internet, which you can verify by turning off Wi-Fi and mobile data.
- Offline
- Once the model is downloaded, everything works in airplane mode. Convenient when traveling or in areas without network coverage.
- Free to use
- No account, no quota, no per-token cost. You only pay for the battery power consumed.
- A limitation to know
- A smartphone realistically tops out at around 4B parameters. For complex reasoning or code, a local LLM on a PC is still far superior.
#What kind of phone you need
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The determining factor is RAM, exactly as on a PC, where VRAM caps the model size. A 3B Q4-quantized model uses about 2 GB, plus the system and application memory. In practice, you need headroom because Android aggressively closes applications that consume too much.
- 6 GB RAM
- The viable minimum. 1B to 3B models in Q4. Close other applications before starting an inference.
- 8 GB RAM
- Comfortable for 3B, possible with 4B. The sweet spot in the latest mid-range segment.
- 12 GB RAM and up
- Recent flagships. Easily handles 4B, and even aggressively quantized 7B, albeit slowly.
- Storage
- Allow 2 to 5 GB of free space per downloaded model. GGUF files are large.
- SoC
- A recent Snapdragon 8-series or equivalent delivers a noticeable speed boost. Entry-level chips work, but remain slow.
#PocketPal: a local LLM in 5 minutes
PocketPal AI is the simplest app to get started. It's an open-source project available on the Play Store that bundles llama.cpp and a catalog of models downloadable directly from Hugging Face, with no command line required.
- 01Install the applicationSearch for « PocketPal AI » in the Google Play Store and install it. The app is free and open source (code available on GitHub).
- 02Open the model catalogIn the Models tab, browse the list of recommended models. PocketPal shows the file size and flags those suited to your RAM.
- 03Download a 1-4B modelChoose a small model to start with, such as a Qwen 3.5 2B (~1.9 GB) or a Granite 4.2 3B (~2.2 GB) with Q4 quantization. The download occurs over the network, only once.
- 04Load and chatClick Load next to the downloaded model, wait for it to load into memory, then open the chat. The first response starts after a few seconds of warm-up.
PocketPal also lets you adjust inference parameters (temperature, context size, number of CPU threads) and import your own GGUF files if you want a model that isn’t in the catalog. For a first experience with a local LLM on Android, it’s the shortest path.
#MLC Chat and GPU acceleration
MLC Chat is the demo application for the MLC LLM project. It compiles models to take advantage of the mobile GPU through the Vulkan/OpenCL API instead of running everything on the CPU. On a recent SoC, this can improve speed and slightly reduce power consumption compared with pure CPU execution.
- 01Retrieve the APKMLC Chat is not always available on the Play Store. Download the official APK from the MLC LLM project website (llm.mlc.ai) and allow installation from an external source.
- 02Download a compiled modelThe app offers models already converted to MLC format (Gemma, Qwen, Granite in smaller sizes). Choose one suited to your RAM.
- 03Start the chatSelect the model, wait for initialization, then start chatting. MLC displays the speed in tokens/second at the bottom of the screen.
#llama.cpp via Termux to go further
For total control—choosing the exact model and quantization and exposing a local server—the advanced path goes through Termux, an Android terminal emulator that provides access to a Linux environment without root. You compile llama.cpp there and run inference from the command line, just as on a Linux PC.
- 01Install and prepare TermuxInstall Termux from F-Droid, open it, then update the core packages.
- 02Install the build toolsGet the compiler, git, and cmake needed to build llama.cpp.
- 03Build llama.cppClone the official repository and compile it. On mobile, the CPU build (ARM NEON) is the most reliable.
- 04Download a GGUF modelDownload a small .gguf file (1-3B in Q4) from Hugging Face with curl or wget.
- 05Run inferenceUse llama-cli to chat, or llama-server to expose an OpenAI-compatible API on localhost that is accessible from a browser.
The server mode is useful: it exposes an OpenAI-compatible endpoint on localhost that you can query from your phone's browser or another app. It's the same approach as the Ollama server on a PC (which listens on port 11434 by default), transferred to your pocket.
#Which 1–4B models actually run
Everything comes down to the 1- to 4-billion-parameter range, in Q4_K_M quantization (the best quality/size compromise, as on PC). Above 4B, a phone slows to a crawl or runs out of RAM. Here are the families that deliver the best results in French.
- Qwen 3.5 2B
- The default small model for 2026 (~1.9 GB in Q4). Excellent for multilingual and French use, with extended context and an Apache 2.0 license. The best starting point for speed on a 6 GB phone.
- Qwen 3.5 4B
- The quality boost (~3.4 GB in Q4) without leaving the mobile tier. Consistent and capable in Q4 if your phone has 8 GB+. Apache 2.0.
- Granite 4.2 3B
- IBM’s small model (~2.2 GB in Q4), highly memory-efficient and token-efficient. A solid offline general-purpose model when RAM is limited. Apache 2.0.
- Gemma 4 E2B
- Google's compact variant (~4.3 GB), multimodal and back under the Apache 2.0 license since April 2026. Best reserved for phones with 8 GB or more.
#Speed, heat, and battery life: what you need to know
LLM inference fully loads the CPU (or GPU), which has three very concrete consequences on a phone: speed remains modest, the device heats up, and the battery drains quickly during generation. Nothing prohibitive, but you need to keep the general scale in mind.
- Speed
- Roughly expect 5 to 15 tokens/second for a 3B Q4 on a mid-to-high-end SoC. It is readable in real time, but nowhere near a PC GPU. A 1B is considerably faster.
- Heat and throttling
- After a long generation, the SoC heats up and throttles, slowing subsequent responses. Let it cool down between long requests.
- Battery
- A sustained session can drain several percent of the battery in a few minutes. For occasional use, it’s negligible; for intensive use, plug in the charger.
- Context
- The larger the context window, the more memory and processing time increase. On mobile, keep the context reasonable (2k-4k tokens) to stay responsive.
#Troubleshooting
- The app closes while loading the model
- Not enough RAM. Close all other applications, choose a smaller model (1B instead of 3B), or use a lighter quantization.
- Very slow responses
- Reduce the context size, lower the number of generated tokens, or switch to a smaller model. Also check that the phone is not throttling because of heat.
- Termux: command not found after compilation
- The build failed silently. Re-run the compilation and read the errors. Make sure git, cmake, and clang are installed.
- Termux abandoned / outdated packages
- You installed the Play Store version. Uninstall it and reinstall Termux from F-Droid or GitHub.
- Inconsistent responses
- Normal for a very small model on a complex task. Simplify the request, or accept that a 1-3B model has reasoning limitations.
#Go further
Once you are comfortable on mobile, you will probably want a more capable machine at home for demanding tasks. These site guides build on this one:
- Run an LLM locally without a GPU (CPU only)
- The same logic as on mobile, but on PC: models recommended by RAM capacity and tips for accelerating pure-CPU operation.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand why Q4_K_M is the recommended compromise, on both phones and PCs.
- Install Ollama
- For a real LLM server at home, accessible from your phone over the local network.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.