API pricing

DeepInfra Pricing

DeepInfra charges from $0.020 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; its current flagship Qwen3.8 2.4T A95B costs $6.00/1M output. The largest context window is 1.0M tokens (DeepSeek Ai DeepSeek V4 Flash). Prices are USD, current as of August 31, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.019/1M

Mistralai Mistral Nemo Instruct 2407

Cheapest output

$0.020/1M

Meta Llama Llama 3.2 3B Instruct

Longest context

1.0M

DeepSeek Ai DeepSeek V4 Flash

Models priced

197

Avg $3.03/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 197 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Meta Llama Llama 3.2 3B Instruct$0.020$0.020131K
Mistralai Mistral Nemo Instruct 2407$0.019$0.030131K
Meta Llama Meta Llama 3.1 8B Instruct Turbo$0.020$0.040131K
Meta Llama Llama 3.2 11B Vision Instruct$0.049$0.049131K
Meta Llama Meta Llama 3.1 8B Instruct$0.030$0.050131K
Sao10k L3 8B Lunaris V1 Turbo$0.040$0.0508K
Meta Llama Llama Guard 3 8B$0.055$0.055131K
Meta Llama Meta Llama 3 8B Instruct$0.030$0.0608K
Mistralai Mistral Small 24B Instruct 2501$0.050$0.08033K
Gemma 4 E4B IT
VisionReasoningTools
$0.020$0.100131K
Google Gemma 3 4B It$0.050$0.100131K
Google Gemma 4 E4b It$0.020$0.100131K
Qwen Qwen2.5 7B Instruct$0.040$0.10033K
GPT OSS 20B
ReasoningTools
$0.030$0.140131K
Microsoft Phi 4$0.070$0.14016K
OpenAI GPT Oss 20B$0.030$0.140131K
Google Gemma 3 12B It$0.050$0.150131K
Qwen Qwen3.5 9B$0.100$0.150262K
Qwen3.5 9B
VisionReasoningTools
$0.100$0.150262K
Google Gemma 3 27B It$0.080$0.160131K
Nvidia Nvidia Nemotron Nano 9B$0.040$0.160131K
GPT OSS 120B
ReasoningTools
$0.037$0.170131K
OpenAI GPT Oss 120B$0.037$0.170131K
DeepSeek Ai DeepSeek V4 Flash$0.090$0.1801.0M
DeepSeek Ai DeepSeek V4 Flash 0731$0.080$0.1801.0M
DeepSeek V4 Flash
ReasoningToolsCache
$0.090$0.1801.0M
DeepSeek V4 Flash 0731
ReasoningToolsCache
$0.080$0.1801.0M
Inclusionai Ling 3.0 Flash$0.060$0.180131K
Meta Llama Llama Guard 4 12B$0.180$0.180164K
Mistralai Mistral Small 3.2 24B Instruct 2506$0.075$0.200128K
Nemotron 3 Nano 30B A3B
ReasoningToolsCache
$0.050$0.200262K
Nvidia Nemotron 3 Nano 30B A3b$0.050$0.200262K
Nvidia Nemotron Content Safety 3.5$0.200$0.200131K
Nvidia Nvidia Nemotron 3.5 Lightning$0.080$0.200262K
Qwen Qwen3 14B$0.120$0.24041K
DeepSeek Ai DeepSeek R1 Distill Qwen 32B$0.270$0.270131K
Qwen Qwen3 32B$0.080$0.28041K
Qwen3 32B
ReasoningTools
$0.080$0.28041K
Llama 4 Scout 17B
VisionTools
$0.100$0.300328K
Meta Llama Llama 4 Scout 17B 16e Instruct$0.100$0.300328K
Llama 3.3 70B Turbo
Tools
$0.100$0.320131K
Meta Llama Llama 3.3 70B Instruct Turbo$0.100$0.320131K
Gemma 4 26B A4B IT
VisionReasoningTools
$0.070$0.340262K
Google Gemma 4 26B A4b It$0.070$0.340262K
Google Gemma 4 31B It Turbo$0.090$0.340262K
DeepSeek Ai DeepSeek V3.2$0.260$0.380164K
DeepSeek-V3.2
ReasoningToolsCache
$0.260$0.380164K
Gemma 4 31B IT
VisionReasoningTools
$0.130$0.380262K
Google Gemma 4 31B It$0.130$0.380262K
Bytedance Seed 2.0 Mini$0.100$0.400256K
GLM-4.7-Flash
ReasoningToolsCache
$0.060$0.400203K
Google Gemini 2.0 Flash 001$0.100$0.4001.0M
Gryphe Mythomax L2 13B$0.400$0.4004K
Llama 3.3 Nemotron Super 49B v1.5
ReasoningTools
$0.400$0.400131K
Meta Llama Llama 3.3 70B Instruct$0.230$0.400131K
Meta Llama Meta Llama 3.1 70B Instruct$0.400$0.400131K
Meta Llama Meta Llama 3.1 70B Instruct Turbo$0.400$0.400131K
Mistralai Mixtral 8X7B Instruct V0.1$0.400$0.40033K
Nvidia Llama 3.3 Nemotron Super 49B V1.5$0.100$0.400131K
Nvidia Nvidia Nemotron 3 Super 120B A12b$0.085$0.400262K
Qwen Qwen2.5 72B Instruct$0.360$0.40033K
Qwen Qwq 32B$0.150$0.400131K
Seed 2.0 Mini
VisionReasoningToolsCache
$0.100$0.400256K
Zai Org Glm 4.7 Flash$0.060$0.400203K
Microsoft Wizardlm 2 8X22B$0.480$0.48066K
GLM-5.3-Flash
VisionReasoningToolsCache
$0.150$0.5001.0M
Qwen Qwen3 30B A3b$0.120$0.50041K
Qwen3 30B A3B
ReasoningTools
$0.120$0.50041K
Qwen Qwen3 235B A22b$0.180$0.54041K
Qwen Qwen3 235B A22b Instruct 2507$0.090$0.550262K
Qwen3 235B-A22B Instruct 2507
Tools
$0.090$0.550262K
Hy3
ReasoningToolsCache
$0.140$0.580262K
Tencent Hy3$0.140$0.580262K
DeepSeek Ai DeepSeek R1 Distill Llama 70B$0.200$0.600131K
Nvidia Llama 3.1 Nemotron 70B Instruct$0.600$0.600131K
OpenAI GPT Oss 120B Turbo$0.150$0.600131K
Qwen Qwen2.5 VL 32B Instruct$0.200$0.600128K
Qwen Qwen3 VL 30B A3b Instruct$0.150$0.600262K
Nousresearch Hermes 3 Llama 3.1 70B$0.700$0.700131K
Anthropic Claude 4 Opus$16.50$82.50200K

Frequently asked questions

How much does DeepInfra cost per 1M tokens?

DeepInfra pricing starts at $0.020 per 1 million output tokens (Meta Llama Llama 3.2 3B Instruct) across 197 models, with its current flagship Qwen3.8 2.4T A95B at $6.00. Older premium models in the catalog list higher. Rates current as of August 31, 2026.

What is the cheapest DeepInfra model?

Meta Llama Llama 3.2 3B Instruct is the cheapest DeepInfra model on output tokens at $0.020 per 1M, while Mistralai Mistral Nemo Instruct 2407 is cheapest on input at $0.019 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest DeepInfra context window?

DeepSeek Ai DeepSeek V4 Flash has the largest context window in the DeepInfra catalog at 1.0M input tokens, with up to 1.0M output tokens per response.

Does DeepInfra support prompt caching?

Yes. GLM-4.7-Flash reads cached input at $0.010 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which DeepInfra models support reasoning or vision?

The DeepInfra catalog includes 50 reasoning models, 29 vision models, 61 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual DeepInfra bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator