Advanced 15 minGateway

LiteLLM: a unified local proxy and cloud

If you switch between a local Ollama for sensitive tasks and cloud APIs (OpenAI, Anthropic) for heavy requests, you quickly end up with three SDKs, three key formats, and three ways to handle errors. LiteLLM is a local proxy that speaks the OpenAI API to your application and routes requests behind the scenes to the right backend—local or cloud—with fallback, rate limiting, and cost tracking. One URL in the code, all the logic in a config.yaml.

By Mohamed Meguedmi·Update 2026-09-01·Tested on Windows, macOS, and Linux

#Why use a local LiteLLM cloud proxy

A typical hybrid stack has two problems. First, application code fills up with if provider == 'openai' / elif provider == 'ollama'. Second, the “local vs. cloud” decision is fixed when the code is written: if Ollama goes down, the app goes down; if you want to switch to Claude for a specific task, you have to redeploy.

LiteLLM solves both. On the application side, you talk to a single OpenAI-compatible endpoint (chat/completions, embeddings, streaming). On the infrastructure side, a config.yaml file describes your models: logical alias, backend, API key, fallback priority. You change the route without touching the code.

i
In two words
LiteLLM = an HTTP gateway that accepts OpenAI-compatible requests and translates them for 100+ providers (Ollama, OpenAI, Anthropic, Mistral, Gemini, Azure, Bedrock…). It is written in Python, runs locally, and can be self-hosted.

#How it works

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The proxy exposes port 4000 by default. Your app sends a POST /chat/completions with model: "chat-fr". LiteLLM checks its config.yaml, sees that chat-fr points to ollama/qwen3.5:9b at localhost:11434, makes the request, normalizes the response to the OpenAI format, and returns the result to the app.

App side
A single URL (http://localhost:4000), a single virtual key, and the standard OpenAI SDK are all you need.
On the proxy side
A model_list maps aliases (chat-fr, code-rapide, analyse-doc) to actual backends.
Routing
Multiple backends for the same alias = load balancing, fallback, automatic retry.
Observability
Logs, latencies, cost per request and per virtual key, exportable to Langfuse, Prometheus, or Postgres.

#Prerequisites

Python 3.10+
LiteLLM is a pip package. A clean venv or pipx is all you need.
Ollama running
On http://localhost:11434 with at least one pulled model. See the Ollama installation guide if needed.
Cloud API keys (optional)
OPENAI_API_KEY, ANTHROPIC_API_KEY if you want to route to the cloud as a fallback.
A .env file
To never commit keys in plain text to config.yaml.
→
Cloud isn’t required
LiteLLM is useful even when running 100% locally. If you have two Ollama models (one small and fast, one large and precise), the proxy handles routing between them and failover if one is saturated.

#1. Installation

Installation with proxy extras
pip install 'litellm[proxy]'

The extra proxy bundles FastAPI, uvicorn, and optional dependencies (Postgres, Redis if you want shared rate limiting). That's sufficient for a quick test. For production, use the official Docker image instead.

Docker variant
docker run -d --name litellm \
  -p 4000:4000 \
  -v $(pwd)/config.yaml:/app/config.yaml \
  --env-file .env \
  ghcr.io/berriai/litellm:main-stable \
  --config /app/config.yaml

Verify that it runs:

Health check
curl http://localhost:4000/health/liveliness

#2. A minimal LiteLLM config.yaml

Create config.yaml next to your project. The structure has three sections: model_list (aliases), litellm_settings (global behavior), and general_settings (auth, database).

config.yaml — Ollama only
model_list:
  - model_name: chat-fr
    litellm_params:
      model: ollama/qwen3.5:9b
      api_base: http://localhost:11434

  - model_name: code-rapide
    litellm_params:
      model: ollama/qwen3-coder:30b
      api_base: http://localhost:11434

litellm_settings:
  drop_params: true
  num_retries: 2
  request_timeout: 60

Start the proxy with this config:

Startup
litellm --config config.yaml --port 4000

On the application side, the OpenAI Python SDK talks directly to the proxy. No LiteLLM dependency in the application code:

client.py
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:4000",
    api_key="sk-fake-local",  # le proxy n'exige pas de vraie clé par défaut
)

resp = client.chat.completions.create(
    model="chat-fr",
    messages=[{"role": "user", "content": "Résume la photosynthèse en 3 lignes."}],
)
print(resp.choices[0].message.content)
i
Note on drop_params
drop_params: true tells LiteLLM to silently ignore parameters that a backend does not support (e.g., logprobs on Ollama). Without it, the proxy returns a 400 and breaks the app.

#3. Add OpenAI and Anthropic

Never hardcode keys. Put them in a .env file next to the config:

.env
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...

Then reference the variables with the os.environ syntax in the YAML — LiteLLM substitutes them at startup:

config.yaml — cloud addition
model_list:
  - model_name: chat-fr
    litellm_params:
      model: ollama/qwen3.5:9b
      api_base: http://localhost:11434

  - model_name: chat-gros
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY

  - model_name: analyse-doc
    litellm_params:
      model: anthropic/claude-haiku-4-5-20251001
      api_key: os.environ/ANTHROPIC_API_KEY

At this point you have three logical aliases. The application code chooses chat-fr for private conversations, chat-gros for long requests, and analyse-doc for reading PDFs. No key appears in the code.

!
Local vs. cloud traffic
An alias pointing to ollama/* remains 100% local. As soon as you call chat-gros or analyse-doc, the request leaves your machine for OpenAI or Anthropic. Choose the alias knowingly on the app side—and log it.

#4. Model-based routing and automatic fallback

This is where the proxy delivers its real value. Two mechanisms to know: multiple entries under the same model_name (load balancing) and the fallbacks key (failover on error).

  1. 01
    Multiple inputs, one alias
    You can declare model_name: chat-fr twice—one pointing to a local Ollama, the other to a cloud Mistral. LiteLLM distributes requests according to the strategy (simple-shuffle by default, or usage-based-routing if you want to optimize cost).
  2. 02
    Explicit fallback
    In litellm_settings, declare which alias takes over if the first one returns an error or times out. The fallback automatically retries on the backup backend.
  3. 03
    Health check active
    LiteLLM periodically pings each model. A Ollama that stops responding is marked unhealthy and removed from the pool until it returns — your requests automatically fail over to the cloud.
config.yaml — Ollama fallback → OpenAI
model_list:
  - model_name: chat-fr
    litellm_params:
      model: ollama/qwen3.5:9b
      api_base: http://localhost:11434

  - model_name: chat-fr-cloud
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY

litellm_settings:
  num_retries: 2
  request_timeout: 30
  fallbacks:
    - chat-fr: ["chat-fr-cloud"]
  context_window_fallbacks:
    - chat-fr: ["chat-fr-cloud"]

With this configuration, your app always calls model: "chat-fr". If Ollama is down, times out, or the prompt exceeds the context window you allocated locally (num_ctx reduced to save VRAM), the proxy transparently switches to GPT-4o-mini. The app sees nothing—just a response, perhaps a little slower.

→
Test the fallback
Stop Ollama (sudo systemctl stop ollama on Linux, or Quit in the Windows tray) and rerun a request. You should see the line "Falling back to model chat-fr-cloud" in the LiteLLM logs. If nothing happens, check that num_retries is not set to 0.

#5. Cost tracking and rate limiting

A hybrid stack has a hidden cost: you think you are using local inference, but 30% of requests actually switch to GPT-4o. LiteLLM calculates the cost of each request from an internal pricing table (kept up to date with public pricing schedules).

To persist logs and expose a dashboard, connect a Postgres instance:

general_settings with Postgres
general_settings:
  master_key: sk-litellm-prod-changeme
  database_url: "postgresql://litellm:pass@localhost:5432/litellm"
  store_model_in_db: true

litellm_settings:
  success_callback: ["langfuse"]   # ou prometheus, datadog, etc.
  cache: true

Once Postgres is connected, the admin UI (http://localhost:4000/ui) shows the cost per virtual key, per model, and per user. You can also create virtual keys with capped budgets—handy for giving a team access without risking runaway costs.

Approximate cost for 1 M output tokens (public prices June 2026, to be recalculated with your provider):

Ollama local (Qwen 3.5 9B Q4)
$0 marginal cost—your electricity and GPU depreciation.
OpenAI gpt-4o-mini
Around $0.60 / 1M output tokens, ideal for a low-cost fallback.
Anthropic Claude Haiku 4.5
Around $5 / 1M output tokens, more expensive but offering excellent value for document analysis.
OpenAI gpt-4o
Around $10 / 1M output tokens, best reserved for tasks where 4o-mini isn’t good enough.

For rate limiting, we define RPM (requests per minute) and TPM (tokens per minute) limits per model. LiteLLM queues requests or returns a 429 depending on your config:

Per-model limits
model_list:
  - model_name: chat-gros
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY
      rpm: 60
      tpm: 100000
!
The master_key is not optional
As soon as you expose the proxy beyond localhost (other machines on the LAN, a Docker container), set a strong master_key in general_settings. Otherwise, anyone on the network can use your cloud API keys.

#Troubleshooting

“Model not found” even though the alias exists
Check the YAML indentation. One extra space under litellm_params and the proxy silently ignores the entry. Run litellm --config config.yaml --debug to see the model_list that was actually loaded.
The fallback does not trigger
num_retries must be ≥ 1 and the timeout must be reached. By default, request_timeout is very generous on the Ollama side—lower it to 30 seconds for responsive fallbacks.
401 error on Ollama
ollama/* does not accept an API key. If you set api_key on a Ollama entry, remove it. LiteLLM passes the key through as-is, which causes it to fail.
False or zero costs
The pricing table depends on the LiteLLM version. Update it (pip install -U 'litellm[proxy]'). For a custom model that isn't listed, manually declare input_cost_per_token and output_cost_per_token in litellm_params.
Abnormal latency on Ollama
The proxy runs a health check every minute. If Ollama is slow to load a model (cold start), the check times out and marks the model unhealthy. Increase health_check_interval or preload the models with ollama run X --keepalive 60m.

#Go further

With this setup, you have a single entry point for all your AI, local or cloud, with a seamless switch between them. The natural next steps:

Frequently asked questions
What is an LLM gateway?+
An LLM gateway (or LLM proxy) is a single gateway between your applications and multiple model providers: your code uses one API format, and the gateway then routes requests to Ollama locally, OpenAI, Anthropic, or any other backend—with key management, fallbacks, and cost tracking in one place. LiteLLM is the most widely used open-source LLM gateway for this purpose.
Is LiteLLM the only possible LLM gateway?+
No: OpenRouter plays a similar role on the cloud-hosted side, and enterprise solutions exist. But for a local, open-source, self-hosted gateway—the one that keeps your keys and logs with you—LiteLLM remains the standard, and that is what this guide covers.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.