DeepInfra Pricing
DeepInfra charges from $0.020 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; its current flagship Kimi K3 costs $14.25/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Gemma 4 E4B IT
Cheapest output
$0.020/1M
Meta Llama Llama 3.2 3B Instruct
Longest context
1.0M
DeepSeek V4 Flash
Models priced
118
Avg $2.28/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 118 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Meta Llama Llama 3.2 3B Instruct | $0.020 | $0.020 | 131K |
| Meta Llama Meta Llama 3.1 8B Instruct Turbo | $0.020 | $0.030 | 131K |
| Mistralai Mistral Nemo Instruct 2407 | $0.020 | $0.040 | 131K |
| Meta Llama Llama 3.2 11B Vision Instruct | $0.049 | $0.049 | 131K |
| Meta Llama Meta Llama 3.1 8B Instruct | $0.030 | $0.050 | 131K |
| Sao10k L3 8B Lunaris V1 Turbo | $0.040 | $0.050 | 8K |
| Meta Llama Llama Guard 3 8B | $0.055 | $0.055 | 131K |
| Meta Llama Meta Llama 3 8B Instruct | $0.030 | $0.060 | 8K |
| Google Gemma 3 4B It | $0.040 | $0.080 | 131K |
| Mistralai Mistral Small 24B Instruct 2501 | $0.050 | $0.080 | 33K |
| Gryphe Mythomax L2 13B | $0.080 | $0.090 | 4K |
| Gemma 4 E4B IT VisionReasoningTools | $0.020 | $0.100 | 131K |
| Google Gemma 3 12B It | $0.050 | $0.100 | 131K |
| Qwen Qwen2.5 7B Instruct | $0.040 | $0.100 | 33K |
| GPT OSS 20B ReasoningTools | $0.030 | $0.140 | 131K |
| Microsoft Phi 4 | $0.070 | $0.140 | 16K |
| OpenAI GPT Oss 20B | $0.040 | $0.150 | 131K |
| Qwen3.5 9B VisionReasoningTools | $0.100 | $0.150 | 262K |
| Google Gemma 3 27B It | $0.090 | $0.160 | 131K |
| Nvidia Nvidia Nemotron Nano 9B | $0.040 | $0.160 | 131K |
| GPT OSS 120B ReasoningTools | $0.037 | $0.170 | 131K |
| DeepSeek V4 Flash ReasoningToolsCache | $0.090 | $0.180 | 1.0M |
| DeepSeek V4 Flash 0731 ReasoningToolsCache | $0.080 | $0.180 | 1.0M |
| Meta Llama Llama Guard 4 12B | $0.180 | $0.180 | 164K |
| Mistralai Mistral Small 3.2 24B Instruct 2506 | $0.075 | $0.200 | 128K |
| Nemotron 3 Nano 30B A3B ReasoningToolsCache | $0.050 | $0.200 | 262K |
| Qwen Qwen3 14B | $0.060 | $0.240 | 41K |
| DeepSeek Ai DeepSeek R1 Distill Qwen 32B | $0.270 | $0.270 | 131K |
| Meta Llama Meta Llama 3.1 70B Instruct Turbo | $0.100 | $0.280 | 131K |
| Qwen Qwen3 32B | $0.100 | $0.280 | 41K |
| Qwen3 32B ReasoningTools | $0.080 | $0.280 | 41K |
| Qwen Qwen3 30B A3b | $0.080 | $0.290 | 41K |
| Llama 4 Scout 17B VisionTools | $0.100 | $0.300 | 328K |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.080 | $0.300 | 328K |
| Nousresearch Hermes 3 Llama 3.1 70B | $0.300 | $0.300 | 131K |
| Llama 3.3 70B Turbo Tools | $0.100 | $0.320 | 131K |
| Gemma 4 26B A4B IT VisionReasoningTools | $0.070 | $0.340 | 262K |
| DeepSeek-V3.2 ReasoningToolsCache | $0.260 | $0.380 | 164K |
| Gemma 4 31B IT VisionReasoningTools | $0.130 | $0.380 | 262K |
| Meta Llama Llama 3.3 70B Instruct Turbo | $0.130 | $0.390 | 131K |
| Qwen Qwen2.5 72B Instruct | $0.120 | $0.390 | 33K |
| GLM-4.7-Flash ReasoningToolsCache | $0.060 | $0.400 | 203K |
| Google Gemini 2.0 Flash 001 | $0.100 | $0.400 | 1.0M |
| Llama 3.3 Nemotron Super 49B v1.5 ReasoningTools | $0.400 | $0.400 | 131K |
| Meta Llama Llama 3.3 70B Instruct | $0.230 | $0.400 | 131K |
| Meta Llama Meta Llama 3.1 70B Instruct | $0.400 | $0.400 | 131K |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.400 | $0.400 | 33K |
| Nvidia Llama 3.3 Nemotron Super 49B V1.5 | $0.100 | $0.400 | 131K |
| Qwen Qwq 32B | $0.150 | $0.400 | 131K |
| OpenAI GPT Oss 120B | $0.050 | $0.450 | 131K |
| Microsoft Wizardlm 2 8X22B | $0.480 | $0.480 | 66K |
| Qwen Qwen3 235B A22b | $0.180 | $0.540 | 41K |
| Qwen3 235B-A22B Instruct 2507 Tools | $0.090 | $0.550 | 262K |
| Hy3 ReasoningToolsCache | $0.140 | $0.580 | 262K |
| DeepSeek Ai DeepSeek R1 Distill Llama 70B | $0.200 | $0.600 | 131K |
| Meta Llama Llama 4 Maverick 17B 128e Instruct Fp8 | $0.150 | $0.600 | 1.0M |
| Nvidia Llama 3.1 Nemotron 70B Instruct | $0.600 | $0.600 | 131K |
| Qwen Qwen2.5 VL 32B Instruct | $0.200 | $0.600 | 128K |
| Qwen Qwen3 235B A22b Instruct 2507 | $0.090 | $0.600 | 262K |
| Sao10k L3.1 70B Euryale V2.2 | $0.650 | $0.750 | 131K |
| Sao10k L3.3 70B Euryale V2.3 | $0.650 | $0.750 | 131K |
| Llama 4 Maverick 17B FP8 Vision | $0.200 | $0.800 | 1.0M |
| Nemotron 3 Nano Omni 30B A3B Reasoning VisionReasoningTools | $0.200 | $0.800 | 262K |
| DeepSeek Ai DeepSeek V3 0324 | $0.250 | $0.880 | 164K |
| DeepSeek Ai DeepSeek V3 | $0.380 | $0.890 | 164K |
| DeepSeek-V3 Tools | $0.320 | $0.890 | 164K |
| DeepSeek-V3.1 ReasoningToolsCache | $0.250 | $0.950 | 164K |
| Qwen3.6 35B A3B VisionReasoningTools | $0.100 | $0.950 | 262K |
| DeepSeek Ai DeepSeek V3.1 | $0.270 | $1.00 | 164K |
| DeepSeek Ai DeepSeek V3.1 Terminus | $0.270 | $1.00 | 164K |
| MiniMax-M2.7 ReasoningToolsCache | $0.250 | $1.00 | 197K |
| Nousresearch Hermes 3 Llama 3.1 405B | $1.00 | $1.00 | 131K |
| Qwen 3.5 35B A3B VisionReasoningToolsCache | $0.140 | $1.00 | 262K |
| Qwen3 Coder 480B A35B Instruct Turbo ToolsCache | $0.300 | $1.00 | 262K |
| MiniMax-M3 VisionReasoningToolsCache | $0.280 | $1.10 | 524K |
| Qwen3-Next 80B-A3B Instruct Tools | $0.090 | $1.10 | 262K |
| MiniMax M2.5 ReasoningToolsCache | $0.150 | $1.15 | 197K |
| Step 3.7 Flash VisionReasoningToolsCache | $0.200 | $1.15 | 262K |
| Inkling Small VisionReasoningToolsCache | $0.450 | $1.20 | 524K |
| Anthropic Claude 4 Opus | $16.50 | $82.50 | 200K |
Frequently asked questions
How much does DeepInfra cost per 1M tokens?
DeepInfra pricing starts at $0.020 per 1 million output tokens (Meta Llama Llama 3.2 3B Instruct) across 118 models, with its current flagship Kimi K3 at $14.25. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest DeepInfra model?
Meta Llama Llama 3.2 3B Instruct is the cheapest DeepInfra model on output tokens at $0.020 per 1M, while Gemma 4 E4B IT is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest DeepInfra context window?
DeepSeek V4 Flash has the largest context window in the DeepInfra catalog at 1.0M input tokens, with up to 16K output tokens per response.
Does DeepInfra support prompt caching?
Yes. GLM-4.7-Flash reads cached input at $0.010 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which DeepInfra models support reasoning or vision?
The DeepInfra catalog includes 41 reasoning models, 23 vision models, 50 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual DeepInfra bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator