LocalAI: the complete OpenAI API, 100% self-hosted
LocalAI (the open-source mudler/LocalAI project, MIT-licensed) is a self-hosted inference server that reproduces the OpenAI APIs, and now also those of Anthropic and ElevenLabs, across more than 60 backends (llama.cpp, vLLM, MLX, whisper.cpp, diffusers…). A single Docker instance serves text, embeddings, audio, images, and video, with integrated AI agents (RAG, MCP, tools). Allow about 30 minutes for a first working deployment on a NVIDIA GPU.
LocalAI is an open-source inference server that exposes exactly the same routes as the OpenAI API — but everything runs on your machine. While Ollama focuses on text chat, LocalAI covers text, embeddings, transcription and audio synthesis, and image generation through a single API. This guide shows how to deploy it in Docker, install models from its gallery, and reconnect an existing OpenAI application without changing the code.
#Why LocalAI
LocalAI (the mudler/LocalAI project on GitHub, under the MIT license, created and maintained by Ettore Di Giacinto and the LocalAI team) presents itself as a “drop-in replacement” for the OpenAI API. In practice, your requests to /v1/chat/completions, /v1/embeddings, /v1/audio/transcriptions, /v1/audio/speech, or /v1/images/generations are sent to a server you host instead of OpenAI’s servers. No token leaves your network, there is no usage-based billing, and there is no quota.
LocalAI’s real value isn’t running yet another chat: it’s unifying multiple modalities behind a single compatible endpoint. One instance serves an LLM for text, an embeddings model for your RAG, Whisper for transcription, Stable Diffusion for images, and now video models. For an application that needs several components, this avoids assembling and maintaining three or four separate servers, each with its own API to learn.
- OpenAI-compatible API
- The same paths, the same JSON payloads. Your official SDKs (openai-python, openai-node) work by changing only the base URL.
- Multi-backend
- LocalAI relies on llama.cpp (GGUF), whisper.cpp, diffusers, piper, and others depending on the model. You don't have to install them one by one.
- Multimodal
- Text, embeddings, audio (STT + TTS), and images on the same instance, each on its own OpenAI route.
- 100% local
- Works offline once the models have been downloaded. No inference telemetry, no cloud dependency.
The project has expanded significantly over the versions since its initial description as a simple OpenAI API clone. Its “drop-in” compatibility now also covers the Anthropic and ElevenLabs APIs across each of its backends. More than 60 backends are supported—including llama.cpp, vLLM, SGLang, transformers, whisper.cpp, diffusers, MLX, and MLX-VLM for Apple Silicon, among others—and can be installed on demand from a backend gallery, without having to bundle them all in advance into a single image.
LocalAI also integrates autonomous AI agents with tool use, RAG, and MCP protocol support, along with a multi-user mode featuring API-key authentication, quotas, and role-based access control. Version 4.1.0 (April 2026) added a distributed cluster mode with intelligent routing based on available VRAM and autoscaling; 4.2.0 (May 2026) added speech and facial recognition, speaker diarization, a Ollama-compatible API, and video generation.
#LocalAI or Ollama, depending on the need
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Both run GGUF through llama.cpp and expose an OpenAI-compatible API for text. The difference is one of scope and philosophy. Ollama focuses on simplicity for text (and a little vision) with a streamlined CLI; LocalAI targets broad coverage—multiple modalities, more backends, more settings, integrated agents—at the cost of a significantly more verbose setup.
| Criterion | Ollama | LocalAI |
|---|---|---|
| Getting started | Immediate (“ollama run”) | More verbose, model YAML |
| Terms | Text (and a little vision) | Text, embeddings, audio, images, video |
| Compatible APIs | OpenAI | OpenAI, Anthropic, ElevenLabs |
| Backends | Primarily llama.cpp | 60+ backends (llama.cpp, vLLM, SGLang, MLX…) |
| Agents / MCP | Non-native | Built-in agents with RAG and MCP |
| Multi-utilisateurs | Non-native | API key, quotas, roles |
| Interface ecosystem | Very comprehensive | More restricted |
| A good choice if | Simple, fast text chat | Multiple modalities through a single API |
There’s nothing stopping you from running both on the same machine: Ollama for everyday interactive chat, and LocalAI as a multimodal gateway for applications that need embeddings, audio, or images behind the same API.
#Prerequisites
LocalAI is deployed most cleanly through Docker, with a dedicated image for your hardware. Plan memory requirements according to the models you’re targeting: in the end, VRAM (or RAM in CPU-only mode) determines what you can actually serve.
- Docker
- Recent Docker Engine or Docker Desktop. Docker Compose recommended for reproducible deployment.
- GPU (optional)
- NVIDIA with the NVIDIA Container Toolkit for CUDA 12 or 13 acceleration. LocalAI also accelerates AMD (ROCm), Intel (oneAPI/SYCL), and Apple Silicon (Metal), with Vulkan as a generic fallback when none of these paths apply. Without a GPU, everything runs on the CPU, more slowly.
- VRAM by size (Q4)
- 3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. Allow extra capacity for an embeddings model and/or Whisper if you serve them in parallel.
- GPU reference points
- RTX 3060 12GB (input) or RTX 4070 12GB comfortably run a 7–14B model; RTX 4090 24GB or a Mac M4 Pro with 24–48 GB of unified memory if you want to go larger.
- Disk space
- Each model weighs several GB, sometimes more for images or video. Set aside a dedicated volume so nothing has to be downloaded again every time the container restarts.
#Deploy LocalAI in Docker
- 01Start a test containerThe fastest command starts LocalAI and exposes the API on port 8080. Use the “-gpu-nvidia-cuda-12” image (or “-cuda-13” on the latest drivers) if you have a NVIDIA card; otherwise, use the default CPU image.
- 02Verify that the API respondsOnce the container is ready, the /v1/models route must return the list (empty at first) in OpenAI format. This indicates that the server is correctly connected to port 8080.
- 03Persisting modelsMount a volume at /models (or /build/models depending on the image) so downloaded models survive restarts. Without a volume, everything is downloaded again on every « docker run ».
- 04Switch to Docker ComposeFor long-term use, describe the service in a docker-compose.yml: image, ports, volume, and GPU reservation. Restart everything with “docker compose up -d”.
#Install a model from the gallery
LocalAI provides a gallery of preconfigured models, also available through “local-ai models list” on the command line or at models.localai.io: each entry includes the right backend, prompt template, and default parameters. You can install a model by name through the API without writing YAML by hand.
For full control, you can also define a model manually in a YAML file placed in the /models folder. This file describes the name exposed by the API, the backend, and the weights file to load.
#A single API for text, embeddings, audio, and images
This is where LocalAI stands out. Each modality uses its standard OpenAI route; you only need to have the appropriate model installed for each one. Here are the four most useful building blocks.
#Migrate an OpenAI app without changing the code
Because the routes and payloads are identical, migrating an application means pointing it to a new base URL and replacing the model names. The official SDKs accept a custom base_url: it is the only parameter to change, whether the application targets the OpenAI, Anthropic, or ElevenLabs API.
- Base URL
- Replace the OpenAI endpoint with http://votre-hote:8080/v1. Often, a simple OPENAI_BASE_URL environment variable.
- Model names
- “gpt-4o” → the name of your local model. This is the main adjustment to make in the code or configuration.
- API key
- Optional locally; use any value if the SDK requires one, or configure a real key on the LocalAI side.
- Behavior differences
- A local 7B does not reason like GPT-4. Adjust your prompts and expectations instead of assuming comparable quality.
#Troubleshooting
- The container starts slowly on the first launch
- LocalAI images and the initial model download are large. That’s normal; subsequent launches are fast if the /models volume is persistent.
- “model not found”
- The request’s “model” field must exactly match the “name” in the gallery or YAML. Check with “curl /v1/models”.
- No GPU acceleration
- Make sure to use an image named "-gpu-nvidia-cuda-12" (or "-cuda-13"), install the NVIDIA Container Toolkit, and pass "--gpus all". Enable DEBUG=true to see which backend was actually selected.
- Slow responses or OOM
- The model exceeds your VRAM and spills over into CPU/RAM. Drop down a level (Q4_K_M rather than Q8_0, or a smaller model) or reduce context_size.
- One modality does not respond
- Every route requires its model: no embeddings without an installed embedding model, no /audio without a Whisper model. Install the missing component from the gallery.
#Go further
LocalAI is just one of the open-weight inference servers available. To make an informed choice, compare it with llama-server (llama.cpp's HTTP server) and Ollama's approach, and fine-tune the memory/quality tradeoff of your models with the quantization guide. Then connect an interface or app to it through its OpenAI endpoint.
- llama-server: a local OpenAI API with llama.cpp
- Ollama vs llama.cpp: which one to choose
- Choose your quantization (Q4, Q5, Q8, FP16)
- Local RAG without coding: Open WebUI, AnythingLLM
- Source: LocalAI GitHub repository
- Source: official LocalAI documentation
- Source: LocalAI backend gallery
#FAQ
Is LocalAI free?+
Does LocalAI support APIs other than OpenAI’s?+
What is the difference between LocalAI and Ollama?+
Is LocalAI secure by default if exposed to the Internet?+
Do you need a GPU for LocalAI?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.