◆ AI at Work — Deploy local AI at work: privacy, compliance, costs · $24 · or all kits $49 →

Cost model · updated 2026

The cost of AI API will it blow your budget?

Cloud API providers burn through a huge share of their revenue. Here’s the calculation to determine whether their next price increase will also blow up your budget.

coût_mensuel_2026 = R × (T_in × P_in + (T_out + T_raisonnement) × P_out)
coût_mensuel_2027 = coût_mensuel_2026 × (1 + r)^Δmois

Calculate my API bill and the tipping point

Enter your monthly volume: the tool calculates today's bill, the bill in 12 and 24 months based on the increase scenario, and the payback period for a local setup.

The hidden increase no one announces

Displayed per-token prices have generally fallen since 2023. But two things have changed behind the scenes: the reasoning tokens are billed as output tokens, and can be several times the length of the visible response; the default routing silently routes “difficult” requests to more expensive reasoning models. A review published by TechAhead in April 2026 thus reports OpenAI API prices “up to 4× higher in certain regions” between March 2025 and January 2026, despite catalog price reductions.

A realistic bill in 2026

WorkloadVisible tokens/requestReasoning tokens/requestCost/req (GPT-5)
Customer support classifier120 in / 40 out1800,0021 €
Code review agent (1 file)2,400 in / 600 out4 8000,050 €
Multi-step research agent8,000 in / 1,200 out22 0000,213 €
Long-context legal RAG62,000 in / 1,500 out9 0000,161 €

Calculation at GPT-5's public rate ($1.25 per million input tokens, $10 for output, with reasoning billed as output), converted at €0.88 per $1 (September 2026).

At 50,000 agent executions per month — a modest mid-market deployment — the “research agent” line alone reaches approximately €10,600/month.

Break-even point versus local deployment

The real question isn't “does a local model match the frontier API on MMLU?” but “can it handle 70% of my traffic acceptably, and how much will migration cost?” Three reference deployments, amortized over 24 months, including electricity (€0.20/kWh, 60% utilization). The threshold compares this monthly cost with the output cost of GPT-5 alone (≈ €8.80 per million tokens):

ConfigurationModelHardware costMonthly costProfitable from
RTX 5090 32 GB (added to an existing PC)Qwen 3 32B Q4_K_M≈ 5 000 €≈ 265 €~30M output tok/month
2 × RTX 5090 (without NVLink, split model)Qwen 2.5 72B Q4_K_M≈ 10 000 €≈ 520 €~59M output tok/month
Mac Studio M5 Ultra 96 GBgpt-oss 120B≈ 6 600 €≈ 300 €~34M output tokens/month

Hardware prices recorded in September 2026 (RTX 5090 between ≈ 5 000 and 5 600 € depending on the vendor, Mac Studio M5 Ultra 96 GB at 6 599 €). The RTX 5090 do not have NVLink: two cards share the model over the PCIe bus.

For the example above (50,000 executions, ~1.16 billion output and reasoning tokens per month, ≈ €10,600/month), the break-even point is exceeded by a very wide margin. But a single GPU is not enough to produce this volume: several cards serving batched requests (vLLM) are required, and some traffic will remain on the API.

The hybrid router: the best of both worlds

If 70% of requests run locally, the API bill drops by approximately 70%, minus hardware costs — provided quality remains acceptable for those requests, which you verify by testing on your own traffic.

Verdict

Current monthly API spendRecommendationWhy
< 400 €Stay on the APIMigration engineering outweighs the savings over 24 months
400 - 1 800 €Optimize prompts, cap reasoningPrompt caching, batch processing (-50%), and capped reasoning effort reduce the bill without leaving the API
1 800 - 8 000 €Hybrid: level 1 local + level 3 APIA high-end GPU pays for itself in a few months if most traffic runs locally
> 8 000 €Migrate aggressivelyThe cost trajectory is no longer sustainable beyond 12 months

Frequently asked questions

Will API prices really increase, or keep falling?

List prices per token have generally fallen since 2023. But the cost per useful response can rise because reasoning tokens are billed as output tokens. A review published by TechAhead in April 2026 reports OpenAI API prices “up to 4× higher in some regions” between March 2025 and January 2026. Budget based on your measured cost per request, not on list price alone.

Is a single high-end GPU (32 GB) enough to replace a frontier API?

For a large share of everyday traffic (code, classification, RAG under 32K context), often yes: a quantized 32B model in Q4_K_M runs at several dozen tokens/s on a RTX 5090. The portion that can actually be replaced depends on your tasks and should be validated against your own traffic. For cutting-edge reasoning and very long contexts, a hybrid router that switches to a frontier API is still necessary.

What inflation rate should you use in the model?

There is no official figure: r is a scenario assumption. For illustration, r = 0.06 per month doubles the bill in a year. Also test r = 0 (stable prices) and a negative value, corresponding to the decline in list prices observed since 2023.

Is fine-tuning an alternative to moving to local deployment?

Fine-tuning with a cloud provider lowers the per-token cost of the adjusted model, but often increases total spending because teams use the lower unit cost to justify more calls. It also locks you into the provider. LoRA fine-tuning of a local model (Qwen3 or Llama) delivers lasting savings.

More free tools