LiteLLM: a unified local proxy and cloud
If you switch between a local Ollama for sensitive tasks and cloud APIs (OpenAI, Anthropic) for heavy requests, you quickly end up with three SDKs, three key formats, and three ways to handle errors. LiteLLM is a local proxy that speaks the OpenAI API to your application and routes requests behind the scenes to the right backend—local or cloud—with fallback, rate limiting, and cost tracking. One URL in the code, all the logic in a config.yaml.
#Why use a local LiteLLM cloud proxy
A typical hybrid stack has two problems. First, application code fills up with if provider == 'openai' / elif provider == 'ollama'. Second, the “local vs. cloud” decision is fixed when the code is written: if Ollama goes down, the app goes down; if you want to switch to Claude for a specific task, you have to redeploy.
LiteLLM solves both. On the application side, you talk to a single OpenAI-compatible endpoint (chat/completions, embeddings, streaming). On the infrastructure side, a config.yaml file describes your models: logical alias, backend, API key, fallback priority. You change the route without touching the code.
#How it works
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The proxy exposes port 4000 by default. Your app sends a POST /chat/completions with model: "chat-fr". LiteLLM checks its config.yaml, sees that chat-fr points to ollama/qwen3.5:9b at localhost:11434, makes the request, normalizes the response to the OpenAI format, and returns the result to the app.
- App side
- A single URL (http://localhost:4000), a single virtual key, and the standard OpenAI SDK are all you need.
- On the proxy side
- A model_list maps aliases (chat-fr, code-rapide, analyse-doc) to actual backends.
- Routing
- Multiple backends for the same alias = load balancing, fallback, automatic retry.
- Observability
- Logs, latencies, cost per request and per virtual key, exportable to Langfuse, Prometheus, or Postgres.
#Prerequisites
- Python 3.10+
- LiteLLM is a pip package. A clean venv or pipx is all you need.
- Ollama running
- On http://localhost:11434 with at least one pulled model. See the Ollama installation guide if needed.
- Cloud API keys (optional)
- OPENAI_API_KEY, ANTHROPIC_API_KEY if you want to route to the cloud as a fallback.
- A .env file
- To never commit keys in plain text to config.yaml.
#1. Installation
The extra proxy bundles FastAPI, uvicorn, and optional dependencies (Postgres, Redis if you want shared rate limiting). That's sufficient for a quick test. For production, use the official Docker image instead.
Verify that it runs:
#2. A minimal LiteLLM config.yaml
Create config.yaml next to your project. The structure has three sections: model_list (aliases), litellm_settings (global behavior), and general_settings (auth, database).
Start the proxy with this config:
On the application side, the OpenAI Python SDK talks directly to the proxy. No LiteLLM dependency in the application code:
#3. Add OpenAI and Anthropic
Never hardcode keys. Put them in a .env file next to the config:
Then reference the variables with the os.environ syntax in the YAML — LiteLLM substitutes them at startup:
At this point you have three logical aliases. The application code chooses chat-fr for private conversations, chat-gros for long requests, and analyse-doc for reading PDFs. No key appears in the code.
#4. Model-based routing and automatic fallback
This is where the proxy delivers its real value. Two mechanisms to know: multiple entries under the same model_name (load balancing) and the fallbacks key (failover on error).
- 01Multiple inputs, one aliasYou can declare model_name: chat-fr twice—one pointing to a local Ollama, the other to a cloud Mistral. LiteLLM distributes requests according to the strategy (simple-shuffle by default, or usage-based-routing if you want to optimize cost).
- 02Explicit fallbackIn litellm_settings, declare which alias takes over if the first one returns an error or times out. The fallback automatically retries on the backup backend.
- 03Health check activeLiteLLM periodically pings each model. A Ollama that stops responding is marked unhealthy and removed from the pool until it returns — your requests automatically fail over to the cloud.
With this configuration, your app always calls model: "chat-fr". If Ollama is down, times out, or the prompt exceeds the context window you allocated locally (num_ctx reduced to save VRAM), the proxy transparently switches to GPT-4o-mini. The app sees nothing—just a response, perhaps a little slower.
#5. Cost tracking and rate limiting
A hybrid stack has a hidden cost: you think you are using local inference, but 30% of requests actually switch to GPT-4o. LiteLLM calculates the cost of each request from an internal pricing table (kept up to date with public pricing schedules).
To persist logs and expose a dashboard, connect a Postgres instance:
Once Postgres is connected, the admin UI (http://localhost:4000/ui) shows the cost per virtual key, per model, and per user. You can also create virtual keys with capped budgets—handy for giving a team access without risking runaway costs.
Approximate cost for 1 M output tokens (public prices June 2026, to be recalculated with your provider):
- Ollama local (Qwen 3.5 9B Q4)
- $0 marginal cost—your electricity and GPU depreciation.
- OpenAI gpt-4o-mini
- Around $0.60 / 1M output tokens, ideal for a low-cost fallback.
- Anthropic Claude Haiku 4.5
- Around $5 / 1M output tokens, more expensive but offering excellent value for document analysis.
- OpenAI gpt-4o
- Around $10 / 1M output tokens, best reserved for tasks where 4o-mini isn’t good enough.
For rate limiting, we define RPM (requests per minute) and TPM (tokens per minute) limits per model. LiteLLM queues requests or returns a 429 depending on your config:
#Troubleshooting
- “Model not found” even though the alias exists
- Check the YAML indentation. One extra space under litellm_params and the proxy silently ignores the entry. Run litellm --config config.yaml --debug to see the model_list that was actually loaded.
- The fallback does not trigger
- num_retries must be ≥ 1 and the timeout must be reached. By default, request_timeout is very generous on the Ollama side—lower it to 30 seconds for responsive fallbacks.
- 401 error on Ollama
- ollama/* does not accept an API key. If you set api_key on a Ollama entry, remove it. LiteLLM passes the key through as-is, which causes it to fail.
- False or zero costs
- The pricing table depends on the LiteLLM version. Update it (pip install -U 'litellm[proxy]'). For a custom model that isn't listed, manually declare input_cost_per_token and output_cost_per_token in litellm_params.
- Abnormal latency on Ollama
- The proxy runs a health check every minute. If Ollama is slow to load a model (cold start), the check times out and marks the model unhealthy. Increase health_check_interval or preload the models with ollama run X --keepalive 60m.
#Go further
With this setup, you have a single entry point for all your AI, local or cloud, with a seamless switch between them. The natural next steps:
What is an LLM gateway?+
Is LiteLLM the only possible LLM gateway?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.