API pricing

DeepInfra Pricing

DeepInfra charges from $0.020 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; its current flagship Kimi K3 costs $14.25/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of August 11, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.020/1M

Gemma 4 E4B IT

Cheapest output

$0.020/1M

Meta Llama Llama 3.2 3B Instruct

Longest context

1.0M

DeepSeek V4 Flash

Models priced

118

Avg $2.28/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 118 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Meta Llama Llama 3.2 3B Instruct$0.020$0.020131K
Meta Llama Meta Llama 3.1 8B Instruct Turbo$0.020$0.030131K
Mistralai Mistral Nemo Instruct 2407$0.020$0.040131K
Meta Llama Llama 3.2 11B Vision Instruct$0.049$0.049131K
Meta Llama Meta Llama 3.1 8B Instruct$0.030$0.050131K
Sao10k L3 8B Lunaris V1 Turbo$0.040$0.0508K
Meta Llama Llama Guard 3 8B$0.055$0.055131K
Meta Llama Meta Llama 3 8B Instruct$0.030$0.0608K
Google Gemma 3 4B It$0.040$0.080131K
Mistralai Mistral Small 24B Instruct 2501$0.050$0.08033K
Gryphe Mythomax L2 13B$0.080$0.0904K
Gemma 4 E4B IT
VisionReasoningTools
$0.020$0.100131K
Google Gemma 3 12B It$0.050$0.100131K
Qwen Qwen2.5 7B Instruct$0.040$0.10033K
GPT OSS 20B
ReasoningTools
$0.030$0.140131K
Microsoft Phi 4$0.070$0.14016K
OpenAI GPT Oss 20B$0.040$0.150131K
Qwen3.5 9B
VisionReasoningTools
$0.100$0.150262K
Google Gemma 3 27B It$0.090$0.160131K
Nvidia Nvidia Nemotron Nano 9B$0.040$0.160131K
GPT OSS 120B
ReasoningTools
$0.037$0.170131K
DeepSeek V4 Flash
ReasoningToolsCache
$0.090$0.1801.0M
DeepSeek V4 Flash 0731
ReasoningToolsCache
$0.080$0.1801.0M
Meta Llama Llama Guard 4 12B$0.180$0.180164K
Mistralai Mistral Small 3.2 24B Instruct 2506$0.075$0.200128K
Nemotron 3 Nano 30B A3B
ReasoningToolsCache
$0.050$0.200262K
Qwen Qwen3 14B$0.060$0.24041K
DeepSeek Ai DeepSeek R1 Distill Qwen 32B$0.270$0.270131K
Meta Llama Meta Llama 3.1 70B Instruct Turbo$0.100$0.280131K
Qwen Qwen3 32B$0.100$0.28041K
Qwen3 32B
ReasoningTools
$0.080$0.28041K
Qwen Qwen3 30B A3b$0.080$0.29041K
Llama 4 Scout 17B
VisionTools
$0.100$0.300328K
Meta Llama Llama 4 Scout 17B 16e Instruct$0.080$0.300328K
Nousresearch Hermes 3 Llama 3.1 70B$0.300$0.300131K
Llama 3.3 70B Turbo
Tools
$0.100$0.320131K
Gemma 4 26B A4B IT
VisionReasoningTools
$0.070$0.340262K
DeepSeek-V3.2
ReasoningToolsCache
$0.260$0.380164K
Gemma 4 31B IT
VisionReasoningTools
$0.130$0.380262K
Meta Llama Llama 3.3 70B Instruct Turbo$0.130$0.390131K
Qwen Qwen2.5 72B Instruct$0.120$0.39033K
GLM-4.7-Flash
ReasoningToolsCache
$0.060$0.400203K
Google Gemini 2.0 Flash 001$0.100$0.4001.0M
Llama 3.3 Nemotron Super 49B v1.5
ReasoningTools
$0.400$0.400131K
Meta Llama Llama 3.3 70B Instruct$0.230$0.400131K
Meta Llama Meta Llama 3.1 70B Instruct$0.400$0.400131K
Mistralai Mixtral 8X7B Instruct V0.1$0.400$0.40033K
Nvidia Llama 3.3 Nemotron Super 49B V1.5$0.100$0.400131K
Qwen Qwq 32B$0.150$0.400131K
OpenAI GPT Oss 120B$0.050$0.450131K
Microsoft Wizardlm 2 8X22B$0.480$0.48066K
Qwen Qwen3 235B A22b$0.180$0.54041K
Qwen3 235B-A22B Instruct 2507
Tools
$0.090$0.550262K
Hy3
ReasoningToolsCache
$0.140$0.580262K
DeepSeek Ai DeepSeek R1 Distill Llama 70B$0.200$0.600131K
Meta Llama Llama 4 Maverick 17B 128e Instruct Fp8$0.150$0.6001.0M
Nvidia Llama 3.1 Nemotron 70B Instruct$0.600$0.600131K
Qwen Qwen2.5 VL 32B Instruct$0.200$0.600128K
Qwen Qwen3 235B A22b Instruct 2507$0.090$0.600262K
Sao10k L3.1 70B Euryale V2.2$0.650$0.750131K
Sao10k L3.3 70B Euryale V2.3$0.650$0.750131K
Llama 4 Maverick 17B FP8
Vision
$0.200$0.8001.0M
Nemotron 3 Nano Omni 30B A3B Reasoning
VisionReasoningTools
$0.200$0.800262K
DeepSeek Ai DeepSeek V3 0324$0.250$0.880164K
DeepSeek Ai DeepSeek V3$0.380$0.890164K
DeepSeek-V3
Tools
$0.320$0.890164K
DeepSeek-V3.1
ReasoningToolsCache
$0.250$0.950164K
Qwen3.6 35B A3B
VisionReasoningTools
$0.100$0.950262K
DeepSeek Ai DeepSeek V3.1$0.270$1.00164K
DeepSeek Ai DeepSeek V3.1 Terminus$0.270$1.00164K
MiniMax-M2.7
ReasoningToolsCache
$0.250$1.00197K
Nousresearch Hermes 3 Llama 3.1 405B$1.00$1.00131K
Qwen 3.5 35B A3B
VisionReasoningToolsCache
$0.140$1.00262K
Qwen3 Coder 480B A35B Instruct Turbo
ToolsCache
$0.300$1.00262K
MiniMax-M3
VisionReasoningToolsCache
$0.280$1.10524K
Qwen3-Next 80B-A3B Instruct
Tools
$0.090$1.10262K
MiniMax M2.5
ReasoningToolsCache
$0.150$1.15197K
Step 3.7 Flash
VisionReasoningToolsCache
$0.200$1.15262K
Inkling Small
VisionReasoningToolsCache
$0.450$1.20524K
Anthropic Claude 4 Opus$16.50$82.50200K

Frequently asked questions

How much does DeepInfra cost per 1M tokens?

DeepInfra pricing starts at $0.020 per 1 million output tokens (Meta Llama Llama 3.2 3B Instruct) across 118 models, with its current flagship Kimi K3 at $14.25. Older premium models in the catalog list higher. Rates current as of August 11, 2026.

What is the cheapest DeepInfra model?

Meta Llama Llama 3.2 3B Instruct is the cheapest DeepInfra model on output tokens at $0.020 per 1M, while Gemma 4 E4B IT is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest DeepInfra context window?

DeepSeek V4 Flash has the largest context window in the DeepInfra catalog at 1.0M input tokens, with up to 16K output tokens per response.

Does DeepInfra support prompt caching?

Yes. GLM-4.7-Flash reads cached input at $0.010 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which DeepInfra models support reasoning or vision?

The DeepInfra catalog includes 41 reasoning models, 23 vision models, 50 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual DeepInfra bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator