API pricing

DeepInfra Pricing

DeepInfra charges between $0.020 and $7.50 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; Qwen3.7 Max is the most expensive at $7.50/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.020/1M

Meta Llama Llama 3.2 3B Instruct

Cheapest output

$0.020/1M

Meta Llama Llama 3.2 3B Instruct

Longest context

1.0M

DeepSeek V4 Flash

Models priced

107

Avg $2.25/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 107 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Meta Llama Llama 3.2 3B Instruct$0.020$0.020131K
Meta Llama Meta Llama 3.1 8B Instruct Turbo$0.020$0.030131K
Mistralai Mistral Nemo Instruct 2407$0.020$0.040131K
Meta Llama Llama 3.2 11B Vision Instruct$0.049$0.049131K
Meta Llama Meta Llama 3.1 8B Instruct$0.030$0.050131K
Sao10k L3 8B Lunaris V1 Turbo$0.040$0.0508K
Meta Llama Llama Guard 3 8B$0.055$0.055131K
Meta Llama Meta Llama 3 8B Instruct$0.030$0.0608K
Google Gemma 3 4B It$0.040$0.080131K
Mistralai Mistral Small 24B Instruct 2501$0.050$0.08033K
Gryphe Mythomax L2 13B$0.080$0.0904K
Google Gemma 3 12B It$0.050$0.100131K
Qwen Qwen2.5 7B Instruct$0.040$0.10033K
GPT OSS 20B
ReasoningTools
$0.030$0.140131K
Microsoft Phi 4$0.070$0.14016K
OpenAI GPT Oss 20B$0.040$0.150131K
Qwen3.5 9B
VisionReasoningTools
$0.100$0.150262K
Google Gemma 3 27B It$0.090$0.160131K
Nvidia Nvidia Nemotron Nano 9B$0.040$0.160131K
GPT OSS 120B
ReasoningTools
$0.037$0.170131K
DeepSeek V4 Flash
ReasoningToolsCache
$0.090$0.1801.0M
Meta Llama Llama Guard 4 12B$0.180$0.180164K
Mistralai Mistral Small 3.2 24B Instruct 2506$0.075$0.200128K
Nemotron 3 Nano 30B A3B
ReasoningTools
$0.050$0.200262K
Qwen Qwen3 14B$0.060$0.24041K
DeepSeek Ai DeepSeek R1 Distill Qwen 32B$0.270$0.270131K
Meta Llama Meta Llama 3.1 70B Instruct Turbo$0.100$0.280131K
Qwen Qwen3 32B$0.100$0.28041K
Qwen3 32B
ReasoningTools
$0.080$0.28041K
Qwen Qwen3 30B A3b$0.080$0.29041K
Llama 4 Scout 17B
VisionTools
$0.100$0.300328K
Meta Llama Llama 4 Scout 17B 16e Instruct$0.080$0.300328K
Nousresearch Hermes 3 Llama 3.1 70B$0.300$0.300131K
Llama 3.3 70B Turbo
Tools
$0.100$0.320131K
Gemma 4 26B A4B IT
VisionReasoningTools
$0.070$0.340262K
DeepSeek-V3.2
ReasoningToolsCache
$0.260$0.380164K
Gemma 4 31B IT
VisionReasoningTools
$0.130$0.380262K
Meta Llama Llama 3.3 70B Instruct Turbo$0.130$0.390131K
Qwen Qwen2.5 72B Instruct$0.120$0.39033K
GLM-4.7-Flash
ReasoningToolsCache
$0.060$0.400203K
Google Gemini 2.0 Flash 001$0.100$0.4001.0M
Llama 3.3 Nemotron Super 49B v1.5
ReasoningTools
$0.400$0.400131K
Meta Llama Llama 3.3 70B Instruct$0.230$0.400131K
Meta Llama Meta Llama 3.1 70B Instruct$0.400$0.400131K
Mistralai Mixtral 8X7B Instruct V0.1$0.400$0.40033K
Nvidia Llama 3.3 Nemotron Super 49B V1.5$0.100$0.400131K
Qwen Qwq 32B$0.150$0.400131K
OpenAI GPT Oss 120B$0.050$0.450131K
Microsoft Wizardlm 2 8X22B$0.480$0.48066K
Qwen Qwen3 235B A22b$0.180$0.54041K
DeepSeek Ai DeepSeek R1 Distill Llama 70B$0.200$0.600131K
Meta Llama Llama 4 Maverick 17B 128e Instruct Fp8$0.150$0.6001.0M
Nvidia Llama 3.1 Nemotron 70B Instruct$0.600$0.600131K
Qwen Qwen2.5 VL 32B Instruct$0.200$0.600128K
Qwen Qwen3 235B A22b Instruct 2507$0.090$0.600262K
Sao10k L3.1 70B Euryale V2.2$0.650$0.750131K
Sao10k L3.3 70B Euryale V2.3$0.650$0.750131K
Llama 4 Maverick 17B FP8
Vision
$0.200$0.8001.0M
Nemotron 3 Nano Omni 30B A3B Reasoning
VisionReasoningTools
$0.200$0.800262K
DeepSeek Ai DeepSeek V3 0324$0.250$0.880164K
DeepSeek Ai DeepSeek V3$0.380$0.890164K
Qwen3.6 35B A3B
VisionReasoningTools
$0.150$0.950262K
DeepSeek Ai DeepSeek V3.1$0.270$1.00164K
DeepSeek Ai DeepSeek V3.1 Terminus$0.270$1.00164K
MiniMax-M2.7
ReasoningToolsCache
$0.250$1.00197K
Nousresearch Hermes 3 Llama 3.1 405B$1.00$1.00131K
Qwen 3.5 35B A3B
VisionReasoningToolsCache
$0.140$1.00262K
Qwen3 Coder 480B A35B Instruct Turbo
ToolsCache
$0.300$1.00262K
Qwen3-Next 80B-A3B Instruct
Tools
$0.090$1.10262K
MiniMax M2.5
ReasoningToolsCache
$0.150$1.15197K
MiniMax-M3
VisionReasoningToolsCache
$0.300$1.20524K
Qwen Qwen3 Coder 480B A35b Instruct Turbo$0.290$1.20262K
Qwen Qwen3 Next 80B A3b Instruct$0.140$1.40262K
Qwen Qwen3 Next 80B A3b Thinking$0.140$1.40262K
Allenai Olmocr 7B 0725 Fp8$0.270$1.5016K
Qwen Qwen3 Coder 480B A35b Instruct$0.400$1.60262K
Zai Org Glm 4.5$0.400$1.60131K
GLM-4.7
ReasoningToolsCache
$0.400$1.75203K
GLM-4.6
ReasoningToolsCache
$0.500$2.00203K
Anthropic Claude 4 Opus$16.50$82.50200K

Frequently asked questions

How much does DeepInfra cost per 1M tokens?

DeepInfra pricing ranges from $0.020 to $7.50 per 1 million output tokens across 107 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest DeepInfra model?

Meta Llama Llama 3.2 3B Instruct is the cheapest DeepInfra model on both axes — $0.020 per 1M input tokens and $0.020 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest DeepInfra context window?

DeepSeek V4 Flash has the largest context window in the DeepInfra catalog at 1.0M input tokens, with up to 16K output tokens per response.

Does DeepInfra support prompt caching?

Yes. GLM-4.7-Flash reads cached input at $0.010 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which DeepInfra models support reasoning or vision?

The DeepInfra catalog includes 33 reasoning models, 17 vision models, 39 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual DeepInfra bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator