API pricing

DeepSeek Pricing

DeepSeek charges between $0.100 and $2.19 per 1 million output tokens, depending on the model. DeepSeek R1 0528 Qwen3 8B is the cheapest at $0.100/1M output and $0.020/1M input; DeepSeek R1 is the most expensive at $2.19/1M output. The largest context window is 1.0M tokens (DeepSeek Chat). Prices are USD, current as of July 20, 2026.

DeepSeek prices well below frontier US labs and discounts cache hits heavily, which is why it shows up in cost-reduction comparisons. The reasoning models bill thinking tokens as output, so measure on your own traffic rather than headline rates.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.020/1M

DeepSeek R1 0528 Qwen3 8B

Cheapest output

$0.100/1M

DeepSeek R1 0528 Qwen3 8B

Longest context

1.0M

DeepSeek Chat

Models priced

21

Avg $0.790/1M output

Pricing by model

All prices in USD per 1 million tokens. All 21 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
DeepSeek R1 0528 Qwen3 8B$0.020$0.10033K
R1 Distill Llama 70B$0.030$0.130131K
R1 Distill Qwen 14B$0.150$0.15033K
R1 Distill Qwen 32B$0.270$0.270131K
DeepSeek Chat
ToolsCache
$0.140$0.2801.0M
DeepSeek Coder$0.140$0.280128K
DeepSeek Reasoner
ReasoningToolsCache
$0.140$0.2801.0M
DeepSeek V4 Flash
ReasoningToolsCache
$0.140$0.2801.0M
DeepSeek V3.2$0.280$0.400164K
DeepSeek V3.2 Exp$0.270$0.400164K
DeepSeek V3.1$0.200$0.800164K
DeepSeek V3 0324$0.240$0.840164K
DeepSeek V4 Pro
ReasoningToolsCache
$0.435$0.8701.0M
DeepSeek V3.1 Terminus$0.230$0.900164K
DeepSeek V3.1 Terminus (exacto)$0.270$1.00131K
DeepSeek$0.270$1.1066K
DeepSeek V3$0.300$1.20164K
R1$0.300$1.20164K
R1 0528$0.400$1.75164K
DeepSeek Prover V2$0.500$2.18164K
DeepSeek R1$0.550$2.1966K

Frequently asked questions

How much does DeepSeek cost per 1M tokens?

DeepSeek pricing ranges from $0.100 to $2.19 per 1 million output tokens across 21 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest DeepSeek model?

DeepSeek R1 0528 Qwen3 8B is the cheapest DeepSeek model on both axes — $0.020 per 1M input tokens and $0.100 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest DeepSeek context window?

DeepSeek Chat has the largest context window in the DeepSeek catalog at 1.0M input tokens, with up to 384K output tokens per response.

Does DeepSeek support prompt caching?

Yes. DeepSeek Chat reads cached input at $0.0028 per 1M tokens, against $0.140 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which DeepSeek models support reasoning or vision?

The DeepSeek catalog includes 3 reasoning models, 4 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual DeepSeek bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Head-to-head comparisons

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator