DeepSeek API: key, pricing, and when to switch to local
The DeepSeek API provides access to DeepSeek models from your own code, with token-based billing and a request format compatible with OpenAI’s. This guide shows how to create a key, make your first call, and read the official pricing table without choosing the wrong row. It does not reproduce any prices: amounts change, and only the provider’s page is authoritative. It ends with the criteria that indicate when a local model becomes simpler or less expensive than the API.
#DeepSeek API: the essentials before you start
DeepSeek offers two entry points that shouldn’t be confused. The free chat site is used in a browser. The API is for developers: your program sends a request, DeepSeek servers return a response, and each exchange is deducted from your balance. This second entry point is what this guide covers.
- What it is
- A pay-as-you-go service hosted by DeepSeek. You don't download anything: the model runs at the provider.
- The format
- Compatible with the OpenAI API. Libraries and tools that know how to talk to OpenAI work by changing two settings: the base URL and the key.
- Billing
- Per token, on a prepaid balance. The pricing table distinguishes sent tokens based on whether they are already cached, as well as generated tokens.
- Your data
- Every request leaves your infrastructure and is processed on the vendor's servers. This is the first thing to examine if you handle personal or confidential data.
- L'alternative
- DeepSeek also publishes its model weights. A version adapted to your hardware can run on your machine, with no per-token charges or data uploads.
#Prerequisites
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
- A developer account
- It’s created on the DeepSeek platform at platform.deepseek.com. This is not the same address as the chat site.
- A payment method
- The service operates on a balance that you credit in advance. Without an available balance, calls are refused.
- A tool for calling the API
- curl is enough for a first test. For a real project, use Python 3 with the openai library, or its Node.js equivalent.
- A safe place for the key
- An environment variable on your workstation, a secrets manager in production. Never the source code.
#Create an DeepSeek API key
- 01Create an account on the platformGo to platform.deepseek.com by typing the address yourself, then sign up. For professional use, use a shared team address rather than a personal address: the account holds the balance and keys, so it must survive a colleague's departure.
- 02Add funds to the balanceThe platform's reload section lets you add credit. Start with a small amount: it's more than enough for testing, and it mechanically limits spending if a poorly written loop runs wild.
- 03Generate the keyIn the API keys section, create a new key and give it a name that indicates what it is for (“essais-poste-clara”, “prod-support”). Copy it immediately: as on most platforms, it is displayed in full only when it is created.
- 04Keep the key out of the codePut it in an environment variable. Your program will read it at startup, and it will appear neither in a Git repository nor in a screenshot.
#First call: the OpenAI-compatible format
The API base URL is https://api.deepseek.com. Pass the key in the Authorization header, preceded by the word Bearer. Before sending a question, start by requesting the list of models your key can call: identifiers change from one generation to the next, and this is the only list that is up to date by construction.
The response is a JSON object in which each entry has an id field. This identifier, copied exactly, is what you must put in your requests. Many tutorials use the historical names deepseek-chat and deepseek-reasoner: before reusing them, check that they appear in the returned list, and read the pricing page to see which model each name currently maps to.
In Python, the official OpenAI library does the job. Only two parameters differ from an OpenAI call: the key and the base address.
The last line is the most useful one going forward. The usage object tells you how many tokens you sent (prompt_tokens) and how many the model generated (completion_tokens). The context-cache documentation describes two additional fields, prompt_cache_hit_tokens and prompt_cache_miss_tokens, which separate input tokens that were already cached from those that were not. Display the object returned by your own call: it is the authoritative source, not an example.
#DeepSeek API pricing: read the official pricing table
All prices are on a single documentation page. Open it alongside this guide: the following paragraphs explain what each line means, not what it costs.
The grid is laid out as a table, with one column per model. Prices are expressed per million tokens. For ordinary French text, a token represents slightly less than one word, but the ratio varies by model and content: for counting, rely on the usage object in your responses rather than a conversion rule.
- Input, cache miss
- The normal cost of the tokens you send: system prompt, conversation history, attached documents, question.
- Input, cache reached (cache hit)
- A discounted price applied to the portion of your request that the service processed recently and kept in cache.
- Output (output)
- The price of the tokens generated by the model. Compare this line with the input line: in APIs of this type, it is generally the higher of the two.
- Reasoning tokens
- A model in reasoning mode drafts a chain of thought before its answer. Check the page to see how these tokens are counted: if they're billed as output, a three-line answer can cost as much as a page.
- Context and maximum output
- The same table shows the context length and the maximum response size. These are not prices, but they cap what a request can cost.
- Currency
- Check the displayed currency. If the pricing grid isn't in euros, add your bank's exchange rate and any fees for top-ups.
#The context cache, the primary source of the gap
The cache works by prefix: if the beginning of a request is identical to the beginning of a recent request, that shared portion is billed at the reduced rate. You don't need to enable anything. However, the order in which you build the request determines what you pay for.
- Stability first
- Put first whatever does not change from one call to the next: system instruction, examples, reference document.
- Variable at the end
- The user's question, the date, and a session identifier go last. Inserting a date on the first line is enough to make every request unique, thereby losing the cache benefit.
- Measure rather than assume
- The cache isn't a guarantee. The portion actually hit is shown in the usage object's cache fields. If it stays close to zero even though your requests are similar, you need to rethink how you're constructing your requests.
#Off-peak hours and temporary discounts
An API pricing schedule may offer a reduced rate during certain hours or a launch period. Three checks are essential before factoring it into a budget.
- Is the discount shown on the page today?
- If the official page mentions neither a time window nor a discount, assume there is none. Do not build a budget around a discount mentioned in an old article.
- Which time zone?
- Ranges are generally given in UTC. In mainland France, add one hour in winter and two in summer.
- Can your workload be moved?
- An idle period only benefits workloads that can wait: overnight summaries, document classification, batch generation. An assistant answering customers during the day won't benefit from it.
#Computing, with your numbers
The cost of a call is the sum of three products: uncached input tokens, cached input tokens, and output tokens, each multiplied by its price and then divided by one million. The function below applies this formula to a response’s usage object. All three prices are left at zero: copy them yourself from the official page for the model you are calling.
#Track your usage
You can check the remaining balance on the platform, and the API exposes an endpoint that returns it in JSON. This is useful for triggering an alert before you run out, rather than afterward.
- Log every call
- Record the date, model, and counters from the usage object. Two weeks of logs under real-world conditions are worth more than any estimate: they're the basis for deciding between API and local.
- Cap the output
- The max_tokens parameter limits the length of a response, and therefore its maximum cost. Set it according to the task instead of leaving the default value.
- Monitor history
- In a conversation, the entire history is sent again on every turn. A discussion with fifty exchanges sends its beginning fifty times, even if caching reduces the cost. Summarize or truncate beyond a certain length.
- Promotional credit and reloaded credit
- If your account has complimentary credit in addition to reloaded credit, the pricing page specifies the order in which they are consumed. Also check whether it has an expiration date.
#API or local model: how to decide
There is no universal threshold at which local becomes cheaper, and this guide does not invent one. The result depends on three numbers only you have: your actual token volume, the current pricing, and the cost of the hardware you would buy. The criteria below often let you decide before you even reach for a calculator.
- Privacy
- Personal data, contracts, proprietary code, client files: with the API, this content goes to a third party established outside the European Union, which falls under the GDPR and must be reviewed with your data protection officer. Locally, the issue does not arise. This is often the deciding criterion by itself.
- Volume and consistency
- Light or irregular use favors the API: you pay nothing when you are not using it. Sustained, predictable use favors local: the machine costs the same whether it handles ten requests or ten thousand.
- Required quality
- The API serves the publisher's large models. On a card with 12 to 24 GB of VRAM, you will run significantly smaller models: allow about 9 GB for a 14B and 19 GB for a 32B in Q4_K_M. If your task requires the large model, running locally assumes a different class of hardware.
- Availability
- The API depends on the provider's load and your connection. Local deployment depends on your machine, which you must monitor and troubleshoot yourself.
- Budget predictability
- The API bill follows usage and can be surprising. Local is a fixed cost known in advance: purchase or rental, electricity, and maintenance time.
- Human time
- An API is quick to connect: one key and a few lines of code. A local server requires installation, updates, and someone who knows what to do when it stops responding. That time has a cost and should be included in the comparison.
#The four-step comparison
- 01MeasureRun your use case through the API for two weeks while logging the usage object. You get a real monthly volume, split between uncached input, cached input, and output.
- 02Encrypt the APIApply the official rate card for the day to this volume. That’s your monthly API cost, with its measurement date.
- 03Encrypt local trafficTake the price of the machine capable of running your target model, spread it over the usage period you choose, then add electricity and maintenance time. The guide on GPU server costs explains this calculation in detail.
- 04Check quality before priceSubmit twenty real queries to the local model your hardware can accommodate and compare the responses with those from the API. If the result is not satisfactory, the cost comparison is irrelevant: you are not comparing the same service.
#The same code for both
Switching between them does not require rewriting your application. Ollama, which listens on http://localhost:11434 by default, also exposes an OpenAI-compatible interface under the /v1 path. The code below switches between the DeepSeek API and a local model based on an environment variable.
The deepseek-r1:14b model is a distilled version that fits on a 12 GB card such as a RTX 3060. This is not the model served by the API: expect less accurate answers on difficult tasks. This setup is specifically intended to demonstrate that on your own queries, in step four of the method.
#Troubleshooting: common errors
The DeepSeek documentation has a page listing error codes. The cases below are the ones encountered during startup; when in doubt, the official page takes precedence over this summary.
- 401, authentication refused
- The key is missing, truncated, or revoked. Check that the environment variable is properly defined in the terminal that launches the program and that no space slipped in during copy-paste.
- 402, insufficient balance
- The account has no credit left. Top up through the platform. An alert on the balance endpoint prevents this from happening in production.
- 400 or 422, invalid request
- The JSON body is malformed or a parameter isn't accepted. The most common cause is a model identifier copied from an old tutorial: go back through the /models list.
- 429, too many requests
- You are sending requests faster than the service can accept them. Space out the calls and try again after an increasingly long delay.
- 500 or 503, server error or overload
- The problem is on the provider's side. Try again after waiting a while, and plan to show a clear message or use a fallback model in your application.
- Very slow response
- During periods of high load, a request may wait a long time before it starts responding. Set a maximum client-side timeout and enable stream mode to display the response as it arrives.
- Higher-than-expected bill
- Three usual suspects: reasoning tokens counted as output, a cache that is rarely hit, and the full conversation history sent back on every turn. The usage object log helps distinguish between them.
#Official sources
Prices, the model list, and billing rules change over time. The provider's pages are the authoritative reference, and you should consult them before making any decision involving specific figures.
#Go further
This guide covers the key, how to read the chart, and the decision-making method. For encryption and installation, these site guides take over:
- How much does an LLM GPU server cost?
- Purchase, rental, or API: the cost items to add up for step three of the comparison. https://quelllm.fr/guide/cout-serveur-gpu-llm
- Free LLM APIs: the real comparison
- Free plans, their quotas, and what happens to your data, if your needs fit within a no-cost tier. https://quelllm.fr/guide/api-llm-gratuites-vs-local
- DeepSeek V4 Pro locally
- The hardware required by the family's flagship model when you want to run it yourself. https://quelllm.fr/guide/guide-deepseek-v4-pro
- Local AI vs. ChatGPT
- The same question, cloud or local, posed for conversational use rather than an API. https://quelllm.fr/guide/ia-locale-vs-chatgpt
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.