Advanced 11 minServers

LocalAI: the complete OpenAI API, 100% self-hosted

Direct response

LocalAI (the open-source mudler/LocalAI project, MIT-licensed) is a self-hosted inference server that reproduces the OpenAI APIs, and now also those of Anthropic and ElevenLabs, across more than 60 backends (llama.cpp, vLLM, MLX, whisper.cpp, diffusers…). A single Docker instance serves text, embeddings, audio, images, and video, with integrated AI agents (RAG, MCP, tools). Allow about 30 minutes for a first working deployment on a NVIDIA GPU.

LocalAI is an open-source inference server that exposes exactly the same routes as the OpenAI API — but everything runs on your machine. While Ollama focuses on text chat, LocalAI covers text, embeddings, transcription and audio synthesis, and image generation through a single API. This guide shows how to deploy it in Docker, install models from its gallery, and reconnect an existing OpenAI application without changing the code.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#Why LocalAI

LocalAI (the mudler/LocalAI project on GitHub, under the MIT license, created and maintained by Ettore Di Giacinto and the LocalAI team) presents itself as a “drop-in replacement” for the OpenAI API. In practice, your requests to /v1/chat/completions, /v1/embeddings, /v1/audio/transcriptions, /v1/audio/speech, or /v1/images/generations are sent to a server you host instead of OpenAI’s servers. No token leaves your network, there is no usage-based billing, and there is no quota.

LocalAI’s real value isn’t running yet another chat: it’s unifying multiple modalities behind a single compatible endpoint. One instance serves an LLM for text, an embeddings model for your RAG, Whisper for transcription, Stable Diffusion for images, and now video models. For an application that needs several components, this avoids assembling and maintaining three or four separate servers, each with its own API to learn.

OpenAI-compatible API
The same paths, the same JSON payloads. Your official SDKs (openai-python, openai-node) work by changing only the base URL.
Multi-backend
LocalAI relies on llama.cpp (GGUF), whisper.cpp, diffusers, piper, and others depending on the model. You don't have to install them one by one.
Multimodal
Text, embeddings, audio (STT + TTS), and images on the same instance, each on its own OpenAI route.
100% local
Works offline once the models have been downloaded. No inference telemetry, no cloud dependency.
i
Open-weight, not magic
LocalAI is a server, not a model. Output quality depends entirely on the open-weight models you load into it—and on your VRAM. A 7B GGUF Q4_K_M remains a 7B model, whether you serve it through Ollama, llama.cpp, or LocalAI.

The project has expanded significantly over the versions since its initial description as a simple OpenAI API clone. Its “drop-in” compatibility now also covers the Anthropic and ElevenLabs APIs across each of its backends. More than 60 backends are supported—including llama.cpp, vLLM, SGLang, transformers, whisper.cpp, diffusers, MLX, and MLX-VLM for Apple Silicon, among others—and can be installed on demand from a backend gallery, without having to bundle them all in advance into a single image.

LocalAI also integrates autonomous AI agents with tool use, RAG, and MCP protocol support, along with a multi-user mode featuring API-key authentication, quotas, and role-based access control. Version 4.1.0 (April 2026) added a distributed cluster mode with intelligent routing based on available VRAM and autoscaling; 4.2.0 (May 2026) added speech and facial recognition, speaker diarization, a Ollama-compatible API, and video generation.


#LocalAI or Ollama, depending on the need

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Both run GGUF through llama.cpp and expose an OpenAI-compatible API for text. The difference is one of scope and philosophy. Ollama focuses on simplicity for text (and a little vision) with a streamlined CLI; LocalAI targets broad coverage—multiple modalities, more backends, more settings, integrated agents—at the cost of a significantly more verbose setup.

LocalAI or Ollama, a quick comparison
CriterionOllamaLocalAI
Getting startedImmediate (“ollama run”)More verbose, model YAML
TermsText (and a little vision)Text, embeddings, audio, images, video
Compatible APIsOpenAIOpenAI, Anthropic, ElevenLabs
BackendsPrimarily llama.cpp60+ backends (llama.cpp, vLLM, SGLang, MLX…)
Agents / MCPNon-nativeBuilt-in agents with RAG and MCP
Multi-utilisateursNon-nativeAPI key, quotas, roles
Interface ecosystemVery comprehensiveMore restricted
A good choice ifSimple, fast text chatMultiple modalities through a single API

There’s nothing stopping you from running both on the same machine: Ollama for everyday interactive chat, and LocalAI as a multimodal gateway for applications that need embeddings, audio, or images behind the same API.

→
The right reflex
If your only need is “chatting with a local LLM,” stay with Ollama; it is simpler. Move to LocalAI as soon as “embeddings,” “transcription,” or “image generation” enters the requirements.

#Prerequisites

LocalAI is deployed most cleanly through Docker, with a dedicated image for your hardware. Plan memory requirements according to the models you’re targeting: in the end, VRAM (or RAM in CPU-only mode) determines what you can actually serve.

Docker
Recent Docker Engine or Docker Desktop. Docker Compose recommended for reproducible deployment.
GPU (optional)
NVIDIA with the NVIDIA Container Toolkit for CUDA 12 or 13 acceleration. LocalAI also accelerates AMD (ROCm), Intel (oneAPI/SYCL), and Apple Silicon (Metal), with Vulkan as a generic fallback when none of these paths apply. Without a GPU, everything runs on the CPU, more slowly.
VRAM by size (Q4)
3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. Allow extra capacity for an embeddings model and/or Whisper if you serve them in parallel.
GPU reference points
RTX 3060 12GB (input) or RTX 4070 12GB comfortably run a 7–14B model; RTX 4090 24GB or a Mac M4 Pro with 24–48 GB of unified memory if you want to go larger.
Disk space
Each model weighs several GB, sometimes more for images or video. Set aside a dedicated volume so nothing has to be downloaded again every time the container restarts.

#Deploy LocalAI in Docker

  1. 01
    Start a test container
    The fastest command starts LocalAI and exposes the API on port 8080. Use the “-gpu-nvidia-cuda-12” image (or “-cuda-13” on the latest drivers) if you have a NVIDIA card; otherwise, use the default CPU image.
  2. 02
    Verify that the API responds
    Once the container is ready, the /v1/models route must return the list (empty at first) in OpenAI format. This indicates that the server is correctly connected to port 8080.
  3. 03
    Persisting models
    Mount a volume at /models (or /build/models depending on the image) so downloaded models survive restarts. Without a volume, everything is downloaded again on every « docker run ».
  4. 04
    Switch to Docker Compose
    For long-term use, describe the service in a docker-compose.yml: image, ports, volume, and GPU reservation. Restart everything with “docker compose up -d”.
Terminal — quick start (CPU)
# Lance LocalAI, API OpenAI-compatible sur le port 8080
docker run -p 8080:8080 --name localai \
  -v $PWD/models:/models \
  localai/localai:latest

# Version GPU NVIDIA (CUDA 12) :
# docker run -p 8080:8080 --gpus all \
#   -v $PWD/models:/models \
#   localai/localai:latest-gpu-nvidia-cuda-12

# Version GPU NVIDIA (CUDA 13, plus récente) :
# docker run -p 8080:8080 --gpus all \
#   -v $PWD/models:/models \
#   localai/localai:latest-gpu-nvidia-cuda-13
docker-compose.yml
services:
  localai:
    image: localai/localai:latest-gpu-nvidia-cuda-12
    container_name: localai
    ports:
      - "8080:8080"
    volumes:
      - ./models:/models
    environment:
      - DEBUG=true
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    restart: unless-stopped
Terminal — verify
# La route est identique à celle d'OpenAI
curl http://localhost:8080/v1/models
!
Don't expose it naked to the Internet
By default, LocalAI listens without authentication. If you need to access it remotely, put it behind a reverse proxy (auth + TLS) or a VPN, and enable an API key. An open inference API is free compute for anyone who comes along.

#Install a model from the gallery

LocalAI provides a gallery of preconfigured models, also available through “local-ai models list” on the command line or at models.localai.io: each entry includes the right backend, prompt template, and default parameters. You can install a model by name through the API without writing YAML by hand.

Terminal — install via the API
# Installe un modèle de la galerie (nom d'exemple)
curl http://localhost:8080/models/apply -H "Content-Type: application/json" -d '{
  "id": "localai@qwen2.5-7b-instruct"
}'

# Suivre l'avancement du téléchargement
curl http://localhost:8080/models/jobs

For full control, you can also define a model manually in a YAML file placed in the /models folder. This file describes the name exposed by the API, the backend, and the weights file to load.

models/qwen.yaml
name: qwen2.5-7b
backend: llama-cpp
parameters:
  model: qwen2.5-7b-instruct-q4_k_m.gguf
context_size: 8192
template:
  chat: |
    <|im_start|>system
    {{.SystemPrompt}}<|im_end|>
    {{.Input}}
→
The name = the “model” field
The “name” in your YAML (or gallery entry) is exactly the value to pass in the “model” field of your requests. This replaces “gpt-4o-mini” when you migrate an app.

#A single API for text, embeddings, audio, and images

This is where LocalAI stands out. Each modality uses its standard OpenAI route; you only need to have the appropriate model installed for each one. Here are the four most useful building blocks.

Terminal — chat (text)
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "qwen2.5-7b",
  "messages": [{"role": "user", "content": "Explique le RAG en une phrase."}]
}'
Terminal — embeddings (RAG)
curl http://localhost:8080/v1/embeddings -H "Content-Type: application/json" -d '{
  "model": "bert-embeddings",
  "input": "Texte à vectoriser pour ma base vectorielle"
}'
Terminal — transcription (Whisper)
curl http://localhost:8080/v1/audio/transcriptions \
  -H "Content-Type: multipart/form-data" \
  -F file="@reunion.wav" \
  -F model="whisper-1"
Terminal — image generation
curl http://localhost:8080/v1/images/generations -H "Content-Type: application/json" -d '{
  "model": "stablediffusion",
  "prompt": "un phare breton sous la pluie, aquarelle",
  "size": "512x512"
}'
i
Loading uses VRAM
Serving text + embeddings + Whisper + Stable Diffusion at the same time adds up the memory usage. LocalAI can unload inactive models (idle timeout) to free VRAM, but on a 12 GB card, alternate heavy workloads instead of keeping everything resident.

#Migrate an OpenAI app without changing the code

Because the routes and payloads are identical, migrating an application means pointing it to a new base URL and replacing the model names. The official SDKs accept a custom base_url: it is the only parameter to change, whether the application targets the OpenAI, Anthropic, or ElevenLabs API.

Python — OpenAI SDK to LocalAI
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8080/v1",  # au lieu de l'endpoint OpenAI
    api_key="sk-localai",                 # ignorée si l'auth n'est pas activée
)

resp = client.chat.completions.create(
    model="qwen2.5-7b",                    # au lieu de "gpt-4o-mini"
    messages=[{"role": "user", "content": "Bonjour !"}],
)
print(resp.choices[0].message.content)
Node.js — same principle
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:8080/v1",
  apiKey: "sk-localai",
});

const resp = await client.chat.completions.create({
  model: "qwen2.5-7b",
  messages: [{ role: "user", content: "Bonjour !" }],
});
console.log(resp.choices[0].message.content);
Base URL
Replace the OpenAI endpoint with http://votre-hote:8080/v1. Often, a simple OPENAI_BASE_URL environment variable.
Model names
“gpt-4o” → the name of your local model. This is the main adjustment to make in the code or configuration.
API key
Optional locally; use any value if the SDK requires one, or configure a real key on the LocalAI side.
Behavior differences
A local 7B does not reason like GPT-4. Adjust your prompts and expectations instead of assuming comparable quality.

#Troubleshooting

The container starts slowly on the first launch
LocalAI images and the initial model download are large. That’s normal; subsequent launches are fast if the /models volume is persistent.
“model not found”
The request’s “model” field must exactly match the “name” in the gallery or YAML. Check with “curl /v1/models”.
No GPU acceleration
Make sure to use an image named "-gpu-nvidia-cuda-12" (or "-cuda-13"), install the NVIDIA Container Toolkit, and pass "--gpus all". Enable DEBUG=true to see which backend was actually selected.
Slow responses or OOM
The model exceeds your VRAM and spills over into CPU/RAM. Drop down a level (Q4_K_M rather than Q8_0, or a smaller model) or reduce context_size.
One modality does not respond
Every route requires its model: no embeddings without an installed embedding model, no /audio without a Whisper model. Install the missing component from the gallery.

#Go further

LocalAI is just one of the open-weight inference servers available. To make an informed choice, compare it with llama-server (llama.cpp's HTTP server) and Ollama's approach, and fine-tune the memory/quality tradeoff of your models with the quantization guide. Then connect an interface or app to it through its OpenAI endpoint.


#FAQ

Is LocalAI free?+
Yes, it's open source under the MIT license and self-hosted with no licensing cost, including for professional or commercial use. The only real costs are the hardware that runs the models and the electricity it consumes, as with any local inference server you run yourself.
Does LocalAI support APIs other than OpenAI’s?+
Yes, since its recent versions. Its “drop-in” compatibility now also covers Anthropic and ElevenLabs APIs on every backend, in addition to OpenAI, substantially expanding the number of applications you can reconnect without rewriting their existing client code, including tools originally designed for those specific cloud providers.
What is the difference between LocalAI and Ollama?+
Both expose an OpenAI-compatible API for text via llama.cpp and are easy to get started with. LocalAI goes much further: more than 60 backends, embeddings, audio, images, video, agents with MCP and RAG, and multi-user mode, at the cost of more verbose configuration than Ollama.
Is LocalAI secure by default if exposed to the Internet?+
No, the API listens without mandatory authentication by default after installation. The project now offers API key authentication, quotas, and role-based access control, but you must enable them explicitly; otherwise, always place LocalAI behind a reverse proxy or VPN.
Do you need a GPU for LocalAI?+
No, LocalAI also runs on the CPU alone, simply more slowly for inference. An NVIDIA, AMD, Intel, or Apple Silicon GPU (via Metal) speeds things up significantly; the project supports these four hardware families, plus Vulkan as a generic multi-vendor fallback, and Jetson L4T for embedded NVIDIA.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.