OpenRouter: what it is, how much it costs, and local deployment ?
OpenRouter is an API gateway that provides access to hundreds of AI models (proprietary and open-weight) through a single account and an OpenAI-compatible format. You pay as you go with prepaid credits, with no markup on token prices; only the credit purchase is charged (5.5% by card, minimum $0.80, or 5% in cryptocurrency). “:free” models are available, limited to 50 requests per day without purchased credits, and 1,000 beyond $10.
OpenRouter is an online service that provides access to hundreds of different proprietary and open AI models through a single API and a single account. Instead of opening an account with every provider, you buy credits once and choose the model for each request. As of September 20, 2026, it is the most widely used gateway among developers and coding tools such as Cline and OpenCode. This page explains what the service does, what it costs, where your data goes, and when a model on your own machine remains the better choice.
#What is OpenRouter?
OpenRouter is an API gateway. The company, founded in 2023, does not train models and, for the most part, does not host them itself: it sits between your application and providers, whether labs such as OpenAI, Anthropic, Google, or Mistral, or hosts serving open-weight models such as Qwen, Llama, or DeepSeek. Its value lies in its position: one account, one key, and one bill, regardless of how many providers are actually used behind the scenes.
Its API follows OpenAI's format. An application already written for OpenAI works after changing two lines: the base address and the key. You select the model by a simple name in the form provider/model, letting you switch between them without touching the rest of the code.
The service earns revenue solely from credit purchases, never from the provider’s listed price for each model: its official documentation states that it adds no markup to the inference price itself. This is an important distinction from other gateways that add an invisible markup directly to the token rate, making comparison with the original provider more difficult.
#How it works
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- A single catalog
- Each model has its own entry: price per million input and output tokens, context size, providers serving it, observed throughput, and latency. Routing variants (“:free”, “:batch”, among others) appear as separate entries, with their own prices and limits.
- Routing between providers
- The same open model is often served by multiple providers at different prices and speeds. OpenRouter chooses based on your preferences (price, throughput, latency) and switches to another provider if the first one goes down, without requiring your code to handle that logic itself.
- Fallbacks between models
- You can specify a list of fallback models: if the first is unavailable or refuses the request, the next one takes over automatically, with no manual intervention or visible interruption on the application side.
- Your own key (BYOK)
- If you already have an account with a provider, you can connect your key and keep OpenRouter as your single entry point. A free monthly allowance is available ($25,000 in inference costs at the list rate for Pay-as-you-go, $200,000 for Enterprise); beyond that, a 5% commission on the normal cost applies, deducted from your OpenRouter credits.
#What it costs
The business model is prepaid. You buy credits, and each request is charged at the provider's rate, with no markup on the token price according to the official documentation. OpenRouter earns a commission deducted when credits are purchased. There is no subscription or monthly commitment to cancel.
| Workstation | What you need to know |
|---|---|
| Token pricing | The provider’s value, displayed exactly as it appears on the model page, with no additional margin according to the official documentation. Output generally costs several times more than input. |
| Commission | 5.5% by credit card ($0.80 minimum) or 5% in cryptocurrency, charged once when credits are purchased. No markup is added to the price displayed by the provider, according to the official documentation. |
| Free models | Some open models are available in a free variant, identified by the “:free” suffix, with its own price (zero), context length, and access points. Limit of 50 requests/day without purchased credits, 1,000/day beyond $10 in credits. |
| Refund | Unused credits are refundable within 24 hours of purchase, according to the official FAQ. After that period, they remain usable but are no longer refundable. |
| Reasoning models | Reasoning tokens, invisible on screen, are billed as output at the same rate. The bill can be several times higher than the displayed response length alone would suggest. |
#Use it in five minutes
- 01Create an account and a keyOn openrouter.ai, create an API key and, if the interface offers the option, cap it at a specific amount. A separate key for each tool makes consumption tracking much easier.
- 02Add creditsA few euros are enough for worry-free testing. The free variants work without credits, with a limit of 50 requests per day; once more than $10 in credits has been purchased for the account, that limit increases to 1,000 requests per day.
- 03Point your tool to the APIBase address https://openrouter.ai/api/v1, your key, and the exact model name. Most tools have an “OpenAI compatible” setting or a built-in direct OpenRouter option.
The exact name of each model is listed on its product page in the online catalog. For open models, it is usually the same identifier as on Hugging Face, simply in lowercase.
#Where your data goes
This is the question to ask before connecting a professional tool to it. A request passes through two parties: OpenRouter, then the provider that runs the model. Each has its own policy, and both must be checked separately before considering whether sensitive data can pass through this path.
The dashboard’s Activity tab lets you review your complete usage history by model and request, helping you spot abnormal usage or a misconfigured tool before the bill comes as a surprise. It’s also the first place to check if you dispute a charged amount, before contacting support for a refund or clarification.
- On the OpenRouter side
- The service says it does not retain prompt content by default; a logging option is available and must be deliberately enabled for debugging. Metadata (model, token count, timestamp) is retained for billing and can be viewed in the Activity tab.
- On the provider side
- Policies vary: some hosting providers retain requests for a while, while others use them to train their own models. Each provider profile specifies this, and an account setting lets you globally exclude providers that train on your data instead of checking provider by provider.
- Legal considerations
- Most servers are outside Europe. For personal data, customer files, or confidential code, check what your contract and the GDPR allow before sending anything, even for a simple one-off technical test.
#OpenRouter or a local model?
| Criterion | OpenRouter | Local model |
|---|---|---|
| Getting started | Five minutes, no hardware | One installation, and a suitable graphics card or Mac |
| Model sizes | Everything, including models with several hundred billion parameters | What fits in your memory: up to 30 billion on 24 GB, 120 billion on a 128 GB Mac |
| Cost | Variable in practice, with no natural ceiling | The hardware once, then the electricity |
| Privacy | Your data passes through two third parties | Nothing leaves the machine |
| Availability | Depends on the network and providers | Works offline |
| Throughput limits | Yes, especially on the free tiers | None |
| Proprietary models | Yes | No |
Many of the open models offered on OpenRouter are exactly the ones you can run locally. In the QuelLLM catalog, Qwen 3 14B weighs 9 GB in Q4, gpt-oss 20B 13 GB, and Mistral Small 3.2 24B 14 GB: a 16 GB card runs them without a bill or limits. OpenRouter becomes relevant for models that don't fit locally, or for usage that's too occasional to justify buying hardware.
The economics depend mainly on volume. For occasional use of a few queries per week, the 5.5% fee on a small credit top-up remains negligible in absolute terms. For intensive use—a coding agent running all day or a product with thousands of users—the cumulative pay-per-token cost quickly exceeds that of a graphics card amortized over several years, especially for modest-sized models that personal hardware can handle with ease.
#Use both together
The most rational setup combines both. A local model handles everyday tasks and anything confidential; a cloud gateway takes over on demand for tasks that require a very large model. A proxy such as LiteLLM puts everything behind a single address, with routing rules and a budget.
In this setup, the BYOK approach described above makes perfect sense: instead of paying OpenRouter's commission on every request sent to a provider you already use directly, you connect your own key and pay the BYOK commission only beyond the free monthly allowance. OpenRouter remains the single entry point for code, while you keep your usual provider billing for most of the volume.
- LiteLLM: a unified proxy for local and cloud
- OpenCode + Ollama: a coding agent without an API bill
- Official OpenRouter documentation
- Source: official FAQ on fees and privacy
- Source: details on BYOK fees and free models
#FAQ
Is OpenRouter free?+
Is OpenRouter reliable and safe?+
What's the difference between OpenRouter and Ollama?+
Is OpenRouter more expensive than going directly to the provider?+
Can you use OpenRouter with Cline, OpenCode, or Open WebUI?+
What is BYOK (bring your own key) on OpenRouter?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.