Local LLMs on iPhone and iPad: apps and models that run vraiment
Running a local LLM on an iPhone is possible, and it really works: Apple Silicon chips (A and M series) and the MLX framework let you run 1- to 8-billion-parameter models offline, directly on the device. This guide separates the apps that deliver on their promises, explains what RAM actually changes depending on your iPhone or iPad, and lays out the real performance benchmarks—without the marketing.
#Why run a local LLM on an iPhone
A local LLM on iOS performs all computation on your device: no server, no subscription, and no data leaves your phone. This is the opposite of an app like ChatGPT, which sends every message to the cloud. The reasons to want this are concrete: complete privacy (medical notes, drafts, code), working on a plane or without a network, and zero recurring cost once the model is downloaded.
The tradeoff is real: an iPhone has neither the memory nor the power of a PC equipped with a RTX 4070. You cannot run a 70B model in your pocket. But a properly quantized 3B model is more than enough to summarize text, rewrite an email, answer simple questions, or translate—all offline and instantly accessible.
#Apple Silicon, MLX, and Neural Engine
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Modern iPhones and iPads share the Apple Silicon architecture with Macs: a system-on-chip (A-series on iPhone, M-series on iPad Pro) and, above all, unified memory shared by the CPU, GPU, and Neural Engine. This common RAM pool makes LLM inference on mobile possible—the model is loaded once and accessible to all cores without copying.
MLX is Apple's machine learning framework, designed for this unified memory. Serious iOS apps use it (or llama.cpp compiled with Metal) to leverage the chip's GPU. That's where most of the inference computation happens: the Metal GPU, not the CPU alone.
- Unified memory
- The CPU, GPU, and ANE share the same RAM. This is the #1 limiting factor for the model size you can load.
- Metal GPU
- Runs inference through MLX or llama.cpp. This is what determines your tokens per second.
- Neural Engine (ANE)
- Powerful but specialized. Little used for LLM text generation to date.
#Prerequisites and RAM by device
On iOS, RAM is the only number that really matters, and Apple does not highlight it. Yet it determines whether a model loads or crashes the app. Here are the guidelines for recent devices.
- iPhone 15 / 15 Plus
- 6 GB of RAM. Comfortable for 1B–3B in Q4. A 3B model runs, while a 7B model is borderline and leaves little headroom for the system.
- iPhone 15 Pro / 16 / 16 Pro
- 8 GB. The mobile sweet spot: smooth 3B, 7B–8B playable in Q4 with a moderate context.
- iPhone 17 Pro
- 12 GB (recent Pro range). Real headroom for a comfortable 7B–8B model and longer contexts.
- iPad Pro M4
- 16 GB or more depending on the configuration. The best iOS target: a smooth 8B model, with larger models becoming feasible.
The memory requirements of a quantized Q4 model are in the same ballpark as on a desktop: a 3B model weighs about 2 GB, a 7B model about 5 GB, and an 8B model slightly more. Add the context memory (KV cache), which grows with conversation length. That’s why a long context can crash an app that was starting up just fine.
#iOS Apps That Hold Up
The App Store is full of “AI” apps that are merely fronts for the cloud. For truly local use, you need an app that downloads the model to the device and runs it with MLX or llama.cpp. Here are the reliable options.
- PocketPal AI
- Free and open source. Downloads GGUF models from Hugging Face and runs them through llama.cpp. Simple chat interface, context and temperature settings. The recommended starting point.
- LLM Farm
- Open source and highly configurable. Supports many formats and lets you fine-tune inference. More technical, but ideal for testing different models.
- Enclave / MLX-based apps
- Apps built on MLX, the Apple framework. Good performance on recent Apple Silicon, often privacy-oriented.
- Private LLM
- Paid app that offers optimized quantized models and a turnkey experience. No configuration required; everything is packaged.
#Install and run a model: step by step
The example below uses PocketPal AI, which is free and representative of the general workflow. The logic is the same in other apps.
- 01Install the appDownload PocketPal AI from the App Store. No account or connection to a server is required.
- 02Choose a modelOpen the built-in model library. The app offers models suited to mobile devices (Qwen 3.5 2B, Granite 4.2 3B, Gemma 4 E2B, Qwen 3.5 4B). Start small: a 2B or 3B model in Q4.
- 03DownloadStart the download (a few hundred MB to ~2 GB depending on the size). Use Wi-Fi. The file remains stored locally on the device.
- 04Load the modelSelect the downloaded model to load it into memory. If the app closes here, you are out of RAM: choose a smaller model or close the other apps.
- 05Discuss offlineOpen a chat and ask a question. Turn on airplane mode to verify: everything keeps working, proving that nothing leaves the device.
For those who want the most bare-metal route, it is also possible to compile llama.cpp yourself, but iOS does not allow arbitrary code to run outside a signed app: in practice, using a dedicated app is the only realistic option for the vast majority of users.
#Which models to choose based on RAM
The right model depends first on your device. Here are concrete recommendations, from the most modest to the most capable, all quantized to Q4.
- iPhone 15 / 15 Plus (6 GB)
- Qwen 3.5 2B (1.9 GB) or Granite 4.2 3B (2.2 GB) for smooth performance; Qwen 3.5 4B (3.4 GB) for higher quality. Stick to short contexts.
- iPhone 16 / 16 Pro (8 GB)
- A 3B–4B for smooth daily use. An 8B–9B (Granite 4.2 8B at 5.3 GB, Qwen 3.5 9B at 6.6 GB, 256k ctx and vision) runs in Q4 with a moderate context—slower but significantly more capable.
- iPhone 17 Pro (12 GB)
- 8B–9B is comfortable, with longer contexts, and Qwen 3.5 9B in Q8 (11 GB) for maximum quality in this tier.
- iPad Pro M4 (16 GB+)
- Fluid Qwen 3.5 9B in real-world use; models up to ~12B (Gemma 4 12B, multimodal, 7.6 GB) become testable, with speed remaining reasonable thanks to the M4 GPU.
For French, Gemma 4 and the Qwen 3.5 models perform well for their size. For summarization, rephrasing, and short answers, a 3B is enough. For more serious reasoning or coding, mobile limitations become apparent quickly—that's when it's time to switch to a machine at home.
#Total privacy: nothing leaves the device
This is local deployment's killer argument on mobile. Once the model is downloaded, no network request is needed to chat. Your prompts, documents, and drafts stay on the iPhone. No conversation telemetry, no training on your data, and no risk of a server leak.
- Check in airplane mode
- The simplest test: disconnect all network access and verify that the chat still works. If it works offline, nothing is transmitted.
- Beware of hybrid apps
- Some “AI” apps silently switch to the cloud when the local model fails. Prefer 100% local, open-source apps to eliminate doubt.
- Sensitive data
- Health notes, work documents under an NDA, personal information: local processing is the only way to be certain they never leave your device.
#Troubleshooting and tips
- The app crashes while loading the model
- Not enough RAM. Close all other apps, choose a smaller model (3B instead of 7B), or use a lighter quantization (Q4 instead of Q5/Q8).
- Very slow responses
- Reduce the context length and the number of generated tokens. Also check the temperature: a hot iPhone throttles and slows inference.
- The iPhone heats up and the battery drains
- Inference pushes the GPU to its limit. That's normal during long sessions. Take breaks, avoid generating continuously, and stay plugged in for intensive use.
- The model gives an off-topic answer
- Normal for a very small model on a complex task. Simplify the request or move up one size if your RAM allows.
- Download blocked
- GGUF files are large. Stay on Wi-Fi and avoid putting the app in the background during the download.
#Go further
Mobile is perfect for portable, confidential use, but for heavy workloads you’ll want a real machine. These site guides build on this one:
- Run an LLM locally on Android
- The Android equivalent: PocketPal, MLC Chat, and llama.cpp via Termux, with the same RAM constraints.
- Install Ollama on macOS (Apple Silicon)
- To move from a phone to a Mac: M1/M2/M3/M4 unified memory goes much further, up to 32B models and beyond.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand why Q4_K_M is the recommended compromise on both mobile and PC, and when to move up to Q5/Q8.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.