Calculator for LLM cost — API vs. self-hosted
How much do ChatGPT, Claude, Gemini, or DeepSeek really cost you per month? At what usage level does a RTX 5090 or an M5 Pro Mac mini become cost-effective? Enter your usage and we’ll tell you.
Your monthly usage
Advanced modes
Enable the optimizations that apply to your use case. Not every model supports every mode—see the notes in the table.
Monthly cost via API (8 fournisseurs)
| Model | Vendor | Input | Output | Total/mois | Notes |
|---|---|---|---|---|---|
| Mistral Small 4★ less expensive | Mistral | $1.5 | $1.5 | 2.6 € | EU-hosted, low-cost |
| DeepSeek V4.1 Flash | DeepSeek | $3.0 | $3.0 | 5.3 € | 50% off-peak |
| GPT-5 mini | OpenAI | $2.5 | $5.0 | 6.6 € | Small model, good value |
| Gemini 3.8 Flash | $7.5 | $9.4 | 14.8 € | Latest Flash (rate guaranteed through 31/12/2026) | |
| Claude Haiku 4.5 | Anthropic | $10.0 | $12.5 | 19.8 € | Very fast, multimodal |
| Claude Sonnet 5 | Anthropic | $20.0 | $25.0 | 39.6 € | 1M context, Anthropic's best quality-to-cost ratio |
| Claude Opus 5.5 | Anthropic | $40.0 | $50.0 | 79.2 € | 1M context, cache at 5% of the price |
| GPT-6 Astra | OpenAI | $100 | $125 | 198 € | OpenAI high end (Sept. 2026) |
Self-hosted monthly cost (4 configs)
| Hardware | VRAM | Purchase TTC | Amort./mois | Electricity/month | Total/mois | Capacity tokens/month | Compatible models |
|---|---|---|---|---|---|---|---|
| RTX 5070 Ti 16 GB (in an existing PC)★ less expensive | 16 GB | 1 300 € | 54.2 € | 21.1 € | 75.3 € | 19.4M ~45 tok/s | 14B is comfortable, 24B is tight in Q4 |
| Mac mini M5 Pro 24 GB⚠ undersized | 24 GB | 1 999 € | 83.3 € | 2.8 € | 86.1 € | 8.6M ~20 tok/s | up to 14B (≈ 16 GB usable out of 24) |
| Mac Studio M5 Max 36 GB | 36 GB | 2 999 € | 125 € | 4.1 € | 129 € | 15.1M ~35 tok/s | up to ~32B models in Q4 (≈ 27 GB usable) |
| RTX 5090 32 GB (in an existing PC) | 32 GB | 5 000 € | 208 € | 31.2 € | 240 € | 34.6M ~80 tok/s | up to 32B in Q4-Q6 (no 70B) |
Verdict
Are you making this argument to replace GitHub Copilot or Cursor? Locally, it is €0/month and your code never leaves for the cloud. The guide “Local coding copilot” gives you the complete setup (Ollama + Cline + Aider, ready-to-use configs) so you can get started in 30 minutes—and tells you honestly where local deployment cannot replace the cloud.
View the kit →Local or ChatGPT / Claude subscription?
Compare the cost of your subscriptions with a local machine shared by the team.
Subscriptions: public prices recorded on 28/09/2026 (before taxes, converted at 0,88 € for 1 $ with 20 % VAT added). Machine: purchase price amortized over 24 months + electricity at 0,20 €/kWh. A local 14B to 32B model does not replace the most powerful models for difficult tasks, nor the web research or image generation included in subscriptions. Beyond a few simultaneous users, a more powerful card becomes necessary.
How much does a GPU server cost to run an LLM?
There are two ways to get your own GPU server: buy it or rent it by the hour. For a few hours of use per day, renting is often cheaper the first year; for continuous use, buying pays for itself quickly. And a machine at home does not send any data out.
Buy: monthly cost over 24 months
| Machine | Pricing | Amortization | Electricity (4 h/day) | Total / month | Up to (Q4) |
|---|---|---|---|---|---|
| RTX 5070 Ti 16 GB (existing PC) | 1 300 € | 54 € | 9 € | ≈ 63 € | 14B |
| Mac mini M5 Pro 24 GB | 1 999 € | 83 € | 2 € | ≈ 85 € | 14B |
| Mac Studio M5 Max 36 GB | 2 999 € | 125 € | 3 € | ≈ 128 € | 32B |
| RTX 5090 32 GB (existing PC) | ≈ 5 000 € | 208 € | 16 € | ≈ 224 € | 32B |
Electricity: load power × 4 h × 30 days at €0.20/kWh, with the machine turned off the rest of the time. Prices recorded in September 2026.
Rent: hourly price
| Rented server | GPU memory | Price / hour | 4 hours per day | 24 hours out of 24 |
|---|---|---|---|---|
| Scaleway L4 (Paris) | 24 GB | €0,79 excluding tax | ≈ 95 € | ≈ 577 € |
| RunPod RTX 5090 | 32 GB | 0,61 à 0,87 € | 73 à 104 € | 443 à 636 € |
| RunPod A100 | 80 GB | 1,05 à 1,40 € | 126 à 168 € | 764 à 1 021 € |
| RunPod H100 | 80 GB | 1,75 à 3,07 € | 210 à 369 € | 1 278 à 2 242 € |
Tarifs publics relevés le 28/09/2026 sur scaleway.com et runpod.io (RunPod : fourchette entre offre communautaire et offre sécurisée, dollars convertis à 0,88 €). 4 h par jour = 120 h par mois ; 24 h sur 24 = 730 h.
Which one should you choose?
- A few hours a day : a rented RTX 5090 costs €73–104 per month, less than the depreciation of a purchased card costing ≈ €5,000.
- Continuously : a RTX 5090 rented 24 hours out of 24 costs 443 to 636 € per month; purchased for ≈ 5 000 €, it pays for itself in 9 to 14 months, including electricity (≈ 95 € per month running continuously).
- Sensitive data : a machine at your location transmits nothing; for rentals, prefer a European hosting provider.
Methodology & assumptions
This calculator provides an estimate, not an accounting quote. The assumptions below are intentionally explicit so you can gauge what differs from your actual case.
API pricing
- Public prices recorded on 28/09/2026 in USD/M tokens; sources: developers.openai.com/api/docs/pricing, platform.claude.com/docs/en/about-claude/pricing, ai.google.dev/gemini-api/docs/pricing, api-docs.deepseek.com, mistral.ai/pricing.
- USD→EUR conversion at a rate of 0.88 (September 2026; may fluctuate by ±5% depending on the day).
- Long-context surcharge (OpenAI, Gemini beyond 200k tokens) NO applied — calc uses a single price per model.
- Cache (a model-specific reduced price, from 2% to 10% of the input price) applied only when the "Cache prompts" toggle is enabled. Otherwise ignored.
- Batch mode (50% off OpenAI/Anthropic/Google/Mistral, 24h delay) applies only when the "Batch mode" toggle is enabled.
- DeepSeek off-peak pricing (-50%) applies only when the corresponding toggle is enabled.
Pricing self-hosted
- Purchase price = France-inclusive VAT prices recorded in September 2026 (toggle "Recoverable VAT" to switch to excluding VAT, -16.67%). Graphics cards assume an existing PC.
- Idle/load wattage = COMPLETE system (PC: PSU+CPU+RAM+SSD+GPU; Mac: all-in-one). Sources: TomsHardware/AnandTech reviews, Apple Power Analyzer, and pcpartpicker community measurements.
- Tokens/month capacity = estimated tok/s × active hours × 30 days. Estimated tok/s on a mixed workload (a model that fits comfortably + a model near the VRAM limit).
- NO counted: maintenance (~1h/month ≈ €30–50 freelance equivalent), internet bandwidth, pure depreciation of resold hardware, defects/RMA.
- Toggle “PC allumé H24” = idle (24h - heures_actives) × idleW. Unchecked = PC powered off outside active hours (0W idle).
Important limitations
- Unequal quality : Llama 3.3 70B locally ≠ GPT-5 for complex reasoning / multimodal tasks. The breakeven point assumes functional equivalence, which may overestimate the profitability of self-hosting for some use cases.
- Availability : self-hosted requires the machine to be powered on to process requests. API = always available, instant scaling.
- Latency : self-hosted = ms (local); API = 200–2000 ms depending on vendor + region.
- Data / GDPR : self-hosted = no data sent to a third party. US API (OpenAI/Anthropic/Google) = EU territorial export. EU-hosted Mistral = native GDPR compliance.
- For more: model comparator · GPU configurator · 250-LLM catalog · quelllm MCP server.