API pricing

DeepSeek Pricing

DeepSeek charges from $0.100 per 1 million output tokens, depending on the model. DeepSeek R1 0528 Qwen3 8B is the cheapest at $0.100/1M output and $0.020/1M input; its current flagship DeepSeek R1 costs $2.19/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of August 31, 2026.

DeepSeek prices well below frontier US labs and discounts cache hits heavily, which is why it shows up in cost-reduction comparisons. The reasoning models bill thinking tokens as output, so measure on your own traffic rather than headline rates.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.020/1M

DeepSeek R1 0528 Qwen3 8B

Cheapest output

$0.100/1M

DeepSeek R1 0528 Qwen3 8B

Longest context

1.0M

DeepSeek V4 Flash

Models priced

22

Avg $0.780/1M output

Pricing by model

All prices in USD per 1 million tokens. All 22 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
DeepSeek R1 0528 Qwen3 8B$0.020$0.10033K
R1 Distill Llama 70B$0.030$0.130131K
R1 Distill Qwen 14B$0.150$0.15033K
R1 Distill Qwen 32B$0.270$0.270131K
DeepSeek Coder$0.140$0.280128K
DeepSeek V4 Flash
ReasoningToolsCache
$0.140$0.2801.0M
DeepSeek V4 Flash Vision Exp
VisionReasoningToolsCache
$0.140$0.2801.0M
DeepSeek V3.2$0.280$0.400164K
DeepSeek V3.2 Exp$0.270$0.400164K
DeepSeek Chat$0.280$0.420131K
DeepSeek Reasoner$0.280$0.420131K
DeepSeek V3.1$0.200$0.800164K
DeepSeek V3 0324$0.240$0.840164K
DeepSeek V4 Pro
ReasoningToolsCache
$0.435$0.8701.0M
DeepSeek V3.1 Terminus$0.230$0.900164K
DeepSeek V3.1 Terminus (exacto)$0.270$1.00131K
DeepSeek$0.270$1.1066K
DeepSeek V3$0.300$1.20164K
R1$0.300$1.20164K
R1 0528$0.400$1.75164K
DeepSeek Prover V2$0.500$2.18164K
DeepSeek R1$0.550$2.1966K

Frequently asked questions

How much does DeepSeek cost per 1M tokens?

DeepSeek pricing starts at $0.100 per 1 million output tokens (DeepSeek R1 0528 Qwen3 8B) across 22 models, with its current flagship DeepSeek R1 at $2.19. Older premium models in the catalog list higher. Rates current as of August 31, 2026.

What is the cheapest DeepSeek model?

DeepSeek R1 0528 Qwen3 8B is the cheapest DeepSeek model on both axes — $0.020 per 1M input tokens and $0.100 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest DeepSeek context window?

DeepSeek V4 Flash has the largest context window in the DeepSeek catalog at 1.0M input tokens, with up to 384K output tokens per response.

Does DeepSeek support prompt caching?

Yes. DeepSeek V4 Flash reads cached input at $0.0028 per 1M tokens, against $0.140 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which DeepSeek models support reasoning or vision?

The DeepSeek catalog includes 3 reasoning models, 1 vision models, 3 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual DeepSeek bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Head-to-head comparisons

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator