Tool 02 · LLM cost

Calculator for LLM cost — API vs. self-hosted

How much do ChatGPT, Claude, Gemini, or DeepSeek really cost you per month? At what usage level does a RTX 5090 or an M5 Pro Mac mini become cost-effective? Enter your usage and we’ll tell you.

Your monthly usage

Reference: 10M input + 2.5M output = ~5,000 average Claude conversations (1,500 input tokens + 500 output tokens per conversation). Adjust according to your usage.

Advanced modes

Enable the optimizations that apply to your use case. Not every model supports every mode—see the notes in the table.

Monthly cost via API (8 fournisseurs)

Click the badges to enable/disable models.
ModelVendorInputOutputTotal/moisNotes
Mistral Small 4★ less expensiveMistral$1.5$1.52.6 €EU-hosted, low-cost
DeepSeek V4.1 FlashDeepSeek$3.0$3.05.3 €50% off-peak
GPT-5 miniOpenAI$2.5$5.06.6 €Small model, good value
Gemini 3.8 FlashGoogle$7.5$9.414.8 €Latest Flash (rate guaranteed through 31/12/2026)
Claude Haiku 4.5Anthropic$10.0$12.519.8 €Very fast, multimodal
Claude Sonnet 5Anthropic$20.0$25.039.6 €1M context, Anthropic's best quality-to-cost ratio
Claude Opus 5.5Anthropic$40.0$50.079.2 €1M context, cache at 5% of the price
GPT-6 AstraOpenAI$100$125198 €OpenAI high end (Sept. 2026)

Self-hosted monthly cost (4 configs)

Hardware amortized over 24 months + electricity (0.20 €/kWh). Capacity = estimated tok/s × 4h/day × 30d; warning if insufficient for 12.5M tokens requested.
HardwareVRAMPurchase TTCAmort./moisElectricity/monthTotal/moisCapacity tokens/monthCompatible models
RTX 5070 Ti 16 GB (in an existing PC)★ less expensive16 GB1 300 €54.2 €21.1 €75.3 €19.4M
~45 tok/s
14B is comfortable, 24B is tight in Q4
Mac mini M5 Pro 24 GB⚠ undersized24 GB1 999 €83.3 €2.8 €86.1 €8.6M
~20 tok/s
up to 14B (≈ 16 GB usable out of 24)
Mac Studio M5 Max 36 GB36 GB2 999 €125 €4.1 €129 €15.1M
~35 tok/s
up to ~32B models in Q4 (≈ 27 GB usable)
RTX 5090 32 GB (in an existing PC)32 GB5 000 €208 €31.2 €240 €34.6M
~80 tok/s
up to 32B in Q4-Q6 (no 70B)

Verdict

Customizable comparison — change the options below.
APIs to compare
Mistral Small 4
2.6 €/mois
Cheaper API ★. No hardware to buy, instant scaling.
Hardware to compare
RTX 5070 Ti 16 GB (in an existing PC)
75.3 €/mois (amorti)
Local tokens, private data, latency in ms.
Self-hosted breakeven
Never (at this volume)
At your current volume, Mistral Small 4 (2.6 €/month) costs less than the electricity alone for RTX 5070 Ti 16 GB (in an existing PC). Increase consumption or change the hardware/API above.
If it's for coding
The same calculation, Copilot version

Are you making this argument to replace GitHub Copilot or Cursor? Locally, it is €0/month and your code never leaves for the cloud. The guide “Local coding copilot” gives you the complete setup (Ollama + Cline + Aider, ready-to-use configs) so you can get started in 30 minutes—and tells you honestly where local deployment cannot replace the cloud.

View the kit →

Local or ChatGPT / Claude subscription?

Compare the cost of your subscriptions with a local machine shared by the team.

Subscriptions
63.4 €
per month, including tax
Local machine
85.1 €
per month for 24 months
Refunded in
32.5 months
savings over 3 years: €217

Subscriptions: public prices recorded on 28/09/2026 (before taxes, converted at 0,88 € for 1 $ with 20 % VAT added). Machine: purchase price amortized over 24 months + electricity at 0,20 €/kWh. A local 14B to 32B model does not replace the most powerful models for difficult tasks, nor the web research or image generation included in subscriptions. Beyond a few simultaneous users, a more powerful card becomes necessary.

How much does a GPU server cost to run an LLM?

There are two ways to get your own GPU server: buy it or rent it by the hour. For a few hours of use per day, renting is often cheaper the first year; for continuous use, buying pays for itself quickly. And a machine at home does not send any data out.

Buy: monthly cost over 24 months

MachinePricingAmortizationElectricity (4 h/day)Total / monthUp to (Q4)
RTX 5070 Ti 16 GB (existing PC)1 300 €54 €9 €≈ 63 €14B
Mac mini M5 Pro 24 GB1 999 €83 €2 €≈ 85 €14B
Mac Studio M5 Max 36 GB2 999 €125 €3 €≈ 128 €32B
RTX 5090 32 GB (existing PC)≈ 5 000 €208 €16 €≈ 224 €32B

Electricity: load power × 4 h × 30 days at €0.20/kWh, with the machine turned off the rest of the time. Prices recorded in September 2026.

Rent: hourly price

Rented serverGPU memoryPrice / hour4 hours per day24 hours out of 24
Scaleway L4 (Paris)24 GB€0,79 excluding tax≈ 95 €≈ 577 €
RunPod RTX 509032 GB0,61 à 0,87 €73 à 104 €443 à 636 €
RunPod A10080 GB1,05 à 1,40 €126 à 168 €764 à 1 021 €
RunPod H10080 GB1,75 à 3,07 €210 à 369 €1 278 à 2 242 €

Tarifs publics relevés le 28/09/2026 sur scaleway.com et runpod.io (RunPod : fourchette entre offre communautaire et offre sécurisée, dollars convertis à 0,88 €). 4 h par jour = 120 h par mois ; 24 h sur 24 = 730 h.

Which one should you choose?

  • A few hours a day : a rented RTX 5090 costs €73–104 per month, less than the depreciation of a purchased card costing ≈ €5,000.
  • Continuously : a RTX 5090 rented 24 hours out of 24 costs 443 to 636 € per month; purchased for ≈ 5 000 €, it pays for itself in 9 to 14 months, including electricity (≈ 95 € per month running continuously).
  • Sensitive data : a machine at your location transmits nothing; for rentals, prefer a European hosting provider.

Methodology & assumptions

This calculator provides an estimate, not an accounting quote. The assumptions below are intentionally explicit so you can gauge what differs from your actual case.

API pricing

  • Public prices recorded on 28/09/2026 in USD/M tokens; sources: developers.openai.com/api/docs/pricing, platform.claude.com/docs/en/about-claude/pricing, ai.google.dev/gemini-api/docs/pricing, api-docs.deepseek.com, mistral.ai/pricing.
  • USD→EUR conversion at a rate of 0.88 (September 2026; may fluctuate by ±5% depending on the day).
  • Long-context surcharge (OpenAI, Gemini beyond 200k tokens) NO applied — calc uses a single price per model.
  • Cache (a model-specific reduced price, from 2% to 10% of the input price) applied only when the "Cache prompts" toggle is enabled. Otherwise ignored.
  • Batch mode (50% off OpenAI/Anthropic/Google/Mistral, 24h delay) applies only when the "Batch mode" toggle is enabled.
  • DeepSeek off-peak pricing (-50%) applies only when the corresponding toggle is enabled.

Pricing self-hosted

  • Purchase price = France-inclusive VAT prices recorded in September 2026 (toggle "Recoverable VAT" to switch to excluding VAT, -16.67%). Graphics cards assume an existing PC.
  • Idle/load wattage = COMPLETE system (PC: PSU+CPU+RAM+SSD+GPU; Mac: all-in-one). Sources: TomsHardware/AnandTech reviews, Apple Power Analyzer, and pcpartpicker community measurements.
  • Tokens/month capacity = estimated tok/s × active hours × 30 days. Estimated tok/s on a mixed workload (a model that fits comfortably + a model near the VRAM limit).
  • NO counted: maintenance (~1h/month ≈ €30–50 freelance equivalent), internet bandwidth, pure depreciation of resold hardware, defects/RMA.
  • Toggle “PC allumé H24” = idle (24h - heures_actives) × idleW. Unchecked = PC powered off outside active hours (0W idle).

Important limitations

  • Unequal quality : Llama 3.3 70B locally ≠ GPT-5 for complex reasoning / multimodal tasks. The breakeven point assumes functional equivalence, which may overestimate the profitability of self-hosting for some use cases.
  • Availability : self-hosted requires the machine to be powered on to process requests. API = always available, instant scaling.
  • Latency : self-hosted = ms (local); API = 200–2000 ms depending on vendor + region.
  • Data / GDPR : self-hosted = no data sent to a third party. US API (OpenAI/Anthropic/Google) = EU territorial export. EU-hosted Mistral = native GDPR compliance.
  • For more: model comparator · GPU configurator · 250-LLM catalog · quelllm MCP server.