Intermediate 10 minCosts

DeepSeek API: key, pricing, and when to switch to local

The DeepSeek API provides access to DeepSeek models from your own code, with token-based billing and a request format compatible with OpenAI’s. This guide shows how to create a key, make your first call, and read the official pricing table without choosing the wrong row. It does not reproduce any prices: amounts change, and only the provider’s page is authoritative. It ends with the criteria that indicate when a local model becomes simpler or less expensive than the API.

By Clara M.·Update 2026-10-01·Tested on Windows, macOS, and Linux

#DeepSeek API: the essentials before you start

DeepSeek offers two entry points that shouldn’t be confused. The free chat site is used in a browser. The API is for developers: your program sends a request, DeepSeek servers return a response, and each exchange is deducted from your balance. This second entry point is what this guide covers.

What it is
A pay-as-you-go service hosted by DeepSeek. You don't download anything: the model runs at the provider.
The format
Compatible with the OpenAI API. Libraries and tools that know how to talk to OpenAI work by changing two settings: the base URL and the key.
Billing
Per token, on a prepaid balance. The pricing table distinguishes sent tokens based on whether they are already cached, as well as generated tokens.
Your data
Every request leaves your infrastructure and is processed on the vendor's servers. This is the first thing to examine if you handle personal or confidential data.
L'alternative
DeepSeek also publishes its model weights. A version adapted to your hardware can run on your machine, with no per-token charges or data uploads.
i
Why this guide contains no prices
A price copied into an article becomes wrong the day the publisher changes its pricing, and nothing tells you. Rather than displaying amounts that will age badly, this guide teaches you how to read the official page and do the calculation using the current figures. Be wary of any price table DeepSeek that lists neither its source nor the date it was recorded.

#Prerequisites

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
A developer account
It’s created on the DeepSeek platform at platform.deepseek.com. This is not the same address as the chat site.
A payment method
The service operates on a balance that you credit in advance. Without an available balance, calls are refused.
A tool for calling the API
curl is enough for a first test. For a real project, use Python 3 with the openai library, or its Node.js equivalent.
A safe place for the key
An environment variable on your workstation, a secrets manager in production. Never the source code.

#Create an DeepSeek API key

  1. 01
    Create an account on the platform
    Go to platform.deepseek.com by typing the address yourself, then sign up. For professional use, use a shared team address rather than a personal address: the account holds the balance and keys, so it must survive a colleague's departure.
  2. 02
    Add funds to the balance
    The platform's reload section lets you add credit. Start with a small amount: it's more than enough for testing, and it mechanically limits spending if a poorly written loop runs wild.
  3. 03
    Generate the key
    In the API keys section, create a new key and give it a name that indicates what it is for (“essais-poste-clara”, “prod-support”). Copy it immediately: as on most platforms, it is displayed in full only when it is created.
  4. 04
    Keep the key out of the code
    Put it in an environment variable. Your program will read it at startup, and it will appear neither in a Git repository nor in a screenshot.
Key management page (account required)
https://platform.deepseek.com/api_keys
Terminal (Linux, macOS)
export DEEPSEEK_API_KEY="collez-votre-cle-ici"
PowerShell (Windows)
$env:DEEPSEEK_API_KEY = "collez-votre-cle-ici"
!
A key is a payment method
Anyone who has your key can spend your balance. Create one key per project so you can revoke one without stopping the others, never put it in JavaScript executed by the browser or in a mobile app, and delete it from the platform at the slightest suspicion of a leak.

#First call: the OpenAI-compatible format

The API base URL is https://api.deepseek.com. Pass the key in the Authorization header, preceded by the word Bearer. Before sending a question, start by requesting the list of models your key can call: identifiers change from one generation to the next, and this is the only list that is up to date by construction.

Terminal: list available models
curl https://api.deepseek.com/models \
  -H "Authorization: Bearer $DEEPSEEK_API_KEY"

The response is a JSON object in which each entry has an id field. This identifier, copied exactly, is what you must put in your requests. Many tutorials use the historical names deepseek-chat and deepseek-reasoner: before reusing them, check that they appear in the returned list, and read the pricing page to see which model each name currently maps to.

Terminal: first question
curl https://api.deepseek.com/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $DEEPSEEK_API_KEY" \
  -d '{
    "model": "IDENTIFIANT_DU_MODELE",
    "messages": [
      {"role": "system", "content": "Tu réponds en français, en trois phrases maximum."},
      {"role": "user", "content": "Explique la notion de token pour un modèle de langage."}
    ],
    "stream": false
  }'

In Python, the official OpenAI library does the job. Only two parameters differ from an OpenAI call: the key and the base address.

Terminal
pip install openai
premier_appel.py
import os
from openai import OpenAI

MODELE = "IDENTIFIANT_DU_MODELE"  # un id renvoyé par /models

client = OpenAI(
    api_key=os.environ["DEEPSEEK_API_KEY"],
    base_url="https://api.deepseek.com",
)

reponse = client.chat.completions.create(
    model=MODELE,
    messages=[
        {"role": "system", "content": "Tu réponds en français, en trois phrases maximum."},
        {"role": "user", "content": "Explique la notion de token pour un modèle de langage."},
    ],
)

print(reponse.choices[0].message.content)
print(reponse.usage)  # le décompte qui sert à la facturation

The last line is the most useful one going forward. The usage object tells you how many tokens you sent (prompt_tokens) and how many the model generated (completion_tokens). The context-cache documentation describes two additional fields, prompt_cache_hit_tokens and prompt_cache_miss_tokens, which separate input tokens that were already cached from those that were not. Display the object returned by your own call: it is the authoritative source, not an example.

#DeepSeek API pricing: read the official pricing table

All prices are on a single documentation page. Open it alongside this guide: the following paragraphs explain what each line means, not what it costs.

Official pricing table (Models & Pricing)
https://api-docs.deepseek.com/quick_start/pricing/

The grid is laid out as a table, with one column per model. Prices are expressed per million tokens. For ordinary French text, a token represents slightly less than one word, but the ratio varies by model and content: for counting, rely on the usage object in your responses rather than a conversion rule.

Input, cache miss
The normal cost of the tokens you send: system prompt, conversation history, attached documents, question.
Input, cache reached (cache hit)
A discounted price applied to the portion of your request that the service processed recently and kept in cache.
Output (output)
The price of the tokens generated by the model. Compare this line with the input line: in APIs of this type, it is generally the higher of the two.
Reasoning tokens
A model in reasoning mode drafts a chain of thought before its answer. Check the page to see how these tokens are counted: if they're billed as output, a three-line answer can cost as much as a page.
Context and maximum output
The same table shows the context length and the maximum response size. These are not prices, but they cap what a request can cost.
Currency
Check the displayed currency. If the pricing grid isn't in euros, add your bank's exchange rate and any fees for top-ups.

#The context cache, the primary source of the gap

The cache works by prefix: if the beginning of a request is identical to the beginning of a recent request, that shared portion is billed at the reduced rate. You don't need to enable anything. However, the order in which you build the request determines what you pay for.

Stability first
Put first whatever does not change from one call to the next: system instruction, examples, reference document.
Variable at the end
The user's question, the date, and a session identifier go last. Inserting a date on the first line is enough to make every request unique, thereby losing the cache benefit.
Measure rather than assume
The cache isn't a guarantee. The portion actually hit is shown in the usage object's cache fields. If it stays close to zero even though your requests are similar, you need to rethink how you're constructing your requests.

#Off-peak hours and temporary discounts

An API pricing schedule may offer a reduced rate during certain hours or a launch period. Three checks are essential before factoring it into a budget.

Is the discount shown on the page today?
If the official page mentions neither a time window nor a discount, assume there is none. Do not build a budget around a discount mentioned in an old article.
Which time zone?
Ranges are generally given in UTC. In mainland France, add one hour in winter and two in summer.
Can your workload be moved?
An idle period only benefits workloads that can wait: overnight summaries, document classification, batch generation. An assistant answering customers during the day won't benefit from it.

#Computing, with your numbers

The cost of a call is the sum of three products: uncached input tokens, cached input tokens, and output tokens, each multiplied by its price and then divided by one million. The function below applies this formula to a response’s usage object. All three prices are left at zero: copy them yourself from the official page for the model you are calling.

cout_appel.py
# Prix par million de tokens, à recopier depuis la page officielle
# Relevé le : (notez la date ici)
PRIX_ENTREE_CACHE_MANQUE = 0.0
PRIX_ENTREE_CACHE_ATTEINT = 0.0
PRIX_SORTIE = 0.0

def cout_appel(usage):
    en_cache = getattr(usage, "prompt_cache_hit_tokens", 0) or 0
    hors_cache = usage.prompt_tokens - en_cache
    total = (
        hors_cache * PRIX_ENTREE_CACHE_MANQUE
        + en_cache * PRIX_ENTREE_CACHE_ATTEINT
        + usage.completion_tokens * PRIX_SORTIE
    )
    return total / 1_000_000

# Exemple : print(cout_appel(reponse.usage))
→
Date your record
Record the date next to the three prices in your configuration file, and reread the official page once a month or before each budget decision. A discrepancy between your calculation and the amount actually consumed is the first sign that the pricing grid has changed.

#Track your usage

You can check the remaining balance on the platform, and the API exposes an endpoint that returns it in JSON. This is useful for triggering an alert before you run out, rather than afterward.

Terminal: check the balance
curl https://api.deepseek.com/user/balance \
  -H "Authorization: Bearer $DEEPSEEK_API_KEY"
Log every call
Record the date, model, and counters from the usage object. Two weeks of logs under real-world conditions are worth more than any estimate: they're the basis for deciding between API and local.
Cap the output
The max_tokens parameter limits the length of a response, and therefore its maximum cost. Set it according to the task instead of leaving the default value.
Monitor history
In a conversation, the entire history is sent again on every turn. A discussion with fifty exchanges sends its beginning fifty times, even if caching reduces the cost. Summarize or truncate beyond a certain length.
Promotional credit and reloaded credit
If your account has complimentary credit in addition to reloaded credit, the pricing page specifies the order in which they are consumed. Also check whether it has an expiration date.

#API or local model: how to decide

There is no universal threshold at which local becomes cheaper, and this guide does not invent one. The result depends on three numbers only you have: your actual token volume, the current pricing, and the cost of the hardware you would buy. The criteria below often let you decide before you even reach for a calculator.

Privacy
Personal data, contracts, proprietary code, client files: with the API, this content goes to a third party established outside the European Union, which falls under the GDPR and must be reviewed with your data protection officer. Locally, the issue does not arise. This is often the deciding criterion by itself.
Volume and consistency
Light or irregular use favors the API: you pay nothing when you are not using it. Sustained, predictable use favors local: the machine costs the same whether it handles ten requests or ten thousand.
Required quality
The API serves the publisher's large models. On a card with 12 to 24 GB of VRAM, you will run significantly smaller models: allow about 9 GB for a 14B and 19 GB for a 32B in Q4_K_M. If your task requires the large model, running locally assumes a different class of hardware.
Availability
The API depends on the provider's load and your connection. Local deployment depends on your machine, which you must monitor and troubleshoot yourself.
Budget predictability
The API bill follows usage and can be surprising. Local is a fixed cost known in advance: purchase or rental, electricity, and maintenance time.
Human time
An API is quick to connect: one key and a few lines of code. A local server requires installation, updates, and someone who knows what to do when it stops responding. That time has a cost and should be included in the comparison.

#The four-step comparison

  1. 01
    Measure
    Run your use case through the API for two weeks while logging the usage object. You get a real monthly volume, split between uncached input, cached input, and output.
  2. 02
    Encrypt the API
    Apply the official rate card for the day to this volume. That’s your monthly API cost, with its measurement date.
  3. 03
    Encrypt local traffic
    Take the price of the machine capable of running your target model, spread it over the usage period you choose, then add electricity and maintenance time. The guide on GPU server costs explains this calculation in detail.
  4. 04
    Check quality before price
    Submit twenty real queries to the local model your hardware can accommodate and compare the responses with those from the API. If the result is not satisfactory, the cost comparison is irrelevant: you are not comparing the same service.

#The same code for both

Switching between them does not require rewriting your application. Ollama, which listens on http://localhost:11434 by default, also exposes an OpenAI-compatible interface under the /v1 path. The code below switches between the DeepSeek API and a local model based on an environment variable.

Terminal: prepare the local model
ollama pull deepseek-r1:14b
client_api_ou_local.py
import os
from openai import OpenAI

LOCAL = os.environ.get("LLM_LOCAL") == "1"

if LOCAL:
    client = OpenAI(api_key="ollama", base_url="http://localhost:11434/v1")
    modele = "deepseek-r1:14b"
else:
    client = OpenAI(
        api_key=os.environ["DEEPSEEK_API_KEY"],
        base_url="https://api.deepseek.com",
    )
    modele = "IDENTIFIANT_DU_MODELE"  # un id renvoyé par /models

reponse = client.chat.completions.create(
    model=modele,
    messages=[{"role": "user", "content": "Résume ce texte en deux phrases : ..."}],
)
print(reponse.choices[0].message.content)

The deepseek-r1:14b model is a distilled version that fits on a 12 GB card such as a RTX 3060. This is not the model served by the API: expect less accurate answers on difficult tasks. This setup is specifically intended to demonstrate that on your own queries, in step four of the method.

i
Nothing requires choosing just one side
Since the code is the same, a hybrid setup is possible: local processing for sensitive content and background volume, and the API for spikes or tasks that exceed the local model’s capabilities. The routing rule should then be based on the nature of the data, not the workload: a confidential document should not be sent to the API because the local server is busy.

#Troubleshooting: common errors

The DeepSeek documentation has a page listing error codes. The cases below are the ones encountered during startup; when in doubt, the official page takes precedence over this summary.

401, authentication refused
The key is missing, truncated, or revoked. Check that the environment variable is properly defined in the terminal that launches the program and that no space slipped in during copy-paste.
402, insufficient balance
The account has no credit left. Top up through the platform. An alert on the balance endpoint prevents this from happening in production.
400 or 422, invalid request
The JSON body is malformed or a parameter isn't accepted. The most common cause is a model identifier copied from an old tutorial: go back through the /models list.
429, too many requests
You are sending requests faster than the service can accept them. Space out the calls and try again after an increasingly long delay.
500 or 503, server error or overload
The problem is on the provider's side. Try again after waiting a while, and plan to show a clear message or use a fallback model in your application.
Very slow response
During periods of high load, a request may wait a long time before it starts responding. Set a maximum client-side timeout and enable stream mode to display the response as it arrives.
Higher-than-expected bill
Three usual suspects: reasoning tokens counted as output, a cache that is rarely hit, and the full conversation history sent back on every turn. The usage object log helps distinguish between them.

#Official sources

Prices, the model list, and billing rules change over time. The provider's pages are the authoritative reference, and you should consult them before making any decision involving specific figures.

Models and pricing
https://api-docs.deepseek.com/quick_start/pricing/
API documentation (first call, guides, error codes)
https://api-docs.deepseek.com/
Weights published by DeepSeek on Hugging Face
https://huggingface.co/deepseek-ai

#Go further

This guide covers the key, how to read the chart, and the decision-making method. For encryption and installation, these site guides take over:

How much does an LLM GPU server cost?
Purchase, rental, or API: the cost items to add up for step three of the comparison. https://quelllm.fr/guide/cout-serveur-gpu-llm
Free LLM APIs: the real comparison
Free plans, their quotas, and what happens to your data, if your needs fit within a no-cost tier. https://quelllm.fr/guide/api-llm-gratuites-vs-local
DeepSeek V4 Pro locally
The hardware required by the family's flagship model when you want to run it yourself. https://quelllm.fr/guide/guide-deepseek-v4-pro
Local AI vs. ChatGPT
The same question, cloud or local, posed for conversational use rather than an API. https://quelllm.fr/guide/ia-locale-vs-chatgpt
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.