DeepInfra Pricing
DeepInfra charges from $0.020 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; its current flagship Qwen3.8 2.4T A95B costs $6.00/1M output. The largest context window is 1.0M tokens (DeepSeek Ai DeepSeek V4 Flash). Prices are USD, current as of August 31, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.019/1M
Mistralai Mistral Nemo Instruct 2407
Cheapest output
$0.020/1M
Meta Llama Llama 3.2 3B Instruct
Longest context
1.0M
DeepSeek Ai DeepSeek V4 Flash
Models priced
197
Avg $3.03/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 197 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Meta Llama Llama 3.2 3B Instruct | $0.020 | $0.020 | 131K |
| Mistralai Mistral Nemo Instruct 2407 | $0.019 | $0.030 | 131K |
| Meta Llama Meta Llama 3.1 8B Instruct Turbo | $0.020 | $0.040 | 131K |
| Meta Llama Llama 3.2 11B Vision Instruct | $0.049 | $0.049 | 131K |
| Meta Llama Meta Llama 3.1 8B Instruct | $0.030 | $0.050 | 131K |
| Sao10k L3 8B Lunaris V1 Turbo | $0.040 | $0.050 | 8K |
| Meta Llama Llama Guard 3 8B | $0.055 | $0.055 | 131K |
| Meta Llama Meta Llama 3 8B Instruct | $0.030 | $0.060 | 8K |
| Mistralai Mistral Small 24B Instruct 2501 | $0.050 | $0.080 | 33K |
| Gemma 4 E4B IT VisionReasoningTools | $0.020 | $0.100 | 131K |
| Google Gemma 3 4B It | $0.050 | $0.100 | 131K |
| Google Gemma 4 E4b It | $0.020 | $0.100 | 131K |
| Qwen Qwen2.5 7B Instruct | $0.040 | $0.100 | 33K |
| GPT OSS 20B ReasoningTools | $0.030 | $0.140 | 131K |
| Microsoft Phi 4 | $0.070 | $0.140 | 16K |
| OpenAI GPT Oss 20B | $0.030 | $0.140 | 131K |
| Google Gemma 3 12B It | $0.050 | $0.150 | 131K |
| Qwen Qwen3.5 9B | $0.100 | $0.150 | 262K |
| Qwen3.5 9B VisionReasoningTools | $0.100 | $0.150 | 262K |
| Google Gemma 3 27B It | $0.080 | $0.160 | 131K |
| Nvidia Nvidia Nemotron Nano 9B | $0.040 | $0.160 | 131K |
| GPT OSS 120B ReasoningTools | $0.037 | $0.170 | 131K |
| OpenAI GPT Oss 120B | $0.037 | $0.170 | 131K |
| DeepSeek Ai DeepSeek V4 Flash | $0.090 | $0.180 | 1.0M |
| DeepSeek Ai DeepSeek V4 Flash 0731 | $0.080 | $0.180 | 1.0M |
| DeepSeek V4 Flash ReasoningToolsCache | $0.090 | $0.180 | 1.0M |
| DeepSeek V4 Flash 0731 ReasoningToolsCache | $0.080 | $0.180 | 1.0M |
| Inclusionai Ling 3.0 Flash | $0.060 | $0.180 | 131K |
| Meta Llama Llama Guard 4 12B | $0.180 | $0.180 | 164K |
| Mistralai Mistral Small 3.2 24B Instruct 2506 | $0.075 | $0.200 | 128K |
| Nemotron 3 Nano 30B A3B ReasoningToolsCache | $0.050 | $0.200 | 262K |
| Nvidia Nemotron 3 Nano 30B A3b | $0.050 | $0.200 | 262K |
| Nvidia Nemotron Content Safety 3.5 | $0.200 | $0.200 | 131K |
| Nvidia Nvidia Nemotron 3.5 Lightning | $0.080 | $0.200 | 262K |
| Qwen Qwen3 14B | $0.120 | $0.240 | 41K |
| DeepSeek Ai DeepSeek R1 Distill Qwen 32B | $0.270 | $0.270 | 131K |
| Qwen Qwen3 32B | $0.080 | $0.280 | 41K |
| Qwen3 32B ReasoningTools | $0.080 | $0.280 | 41K |
| Llama 4 Scout 17B VisionTools | $0.100 | $0.300 | 328K |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.100 | $0.300 | 328K |
| Llama 3.3 70B Turbo Tools | $0.100 | $0.320 | 131K |
| Meta Llama Llama 3.3 70B Instruct Turbo | $0.100 | $0.320 | 131K |
| Gemma 4 26B A4B IT VisionReasoningTools | $0.070 | $0.340 | 262K |
| Google Gemma 4 26B A4b It | $0.070 | $0.340 | 262K |
| Google Gemma 4 31B It Turbo | $0.090 | $0.340 | 262K |
| DeepSeek Ai DeepSeek V3.2 | $0.260 | $0.380 | 164K |
| DeepSeek-V3.2 ReasoningToolsCache | $0.260 | $0.380 | 164K |
| Gemma 4 31B IT VisionReasoningTools | $0.130 | $0.380 | 262K |
| Google Gemma 4 31B It | $0.130 | $0.380 | 262K |
| Bytedance Seed 2.0 Mini | $0.100 | $0.400 | 256K |
| GLM-4.7-Flash ReasoningToolsCache | $0.060 | $0.400 | 203K |
| Google Gemini 2.0 Flash 001 | $0.100 | $0.400 | 1.0M |
| Gryphe Mythomax L2 13B | $0.400 | $0.400 | 4K |
| Llama 3.3 Nemotron Super 49B v1.5 ReasoningTools | $0.400 | $0.400 | 131K |
| Meta Llama Llama 3.3 70B Instruct | $0.230 | $0.400 | 131K |
| Meta Llama Meta Llama 3.1 70B Instruct | $0.400 | $0.400 | 131K |
| Meta Llama Meta Llama 3.1 70B Instruct Turbo | $0.400 | $0.400 | 131K |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.400 | $0.400 | 33K |
| Nvidia Llama 3.3 Nemotron Super 49B V1.5 | $0.100 | $0.400 | 131K |
| Nvidia Nvidia Nemotron 3 Super 120B A12b | $0.085 | $0.400 | 262K |
| Qwen Qwen2.5 72B Instruct | $0.360 | $0.400 | 33K |
| Qwen Qwq 32B | $0.150 | $0.400 | 131K |
| Seed 2.0 Mini VisionReasoningToolsCache | $0.100 | $0.400 | 256K |
| Zai Org Glm 4.7 Flash | $0.060 | $0.400 | 203K |
| Microsoft Wizardlm 2 8X22B | $0.480 | $0.480 | 66K |
| GLM-5.3-Flash VisionReasoningToolsCache | $0.150 | $0.500 | 1.0M |
| Qwen Qwen3 30B A3b | $0.120 | $0.500 | 41K |
| Qwen3 30B A3B ReasoningTools | $0.120 | $0.500 | 41K |
| Qwen Qwen3 235B A22b | $0.180 | $0.540 | 41K |
| Qwen Qwen3 235B A22b Instruct 2507 | $0.090 | $0.550 | 262K |
| Qwen3 235B-A22B Instruct 2507 Tools | $0.090 | $0.550 | 262K |
| Hy3 ReasoningToolsCache | $0.140 | $0.580 | 262K |
| Tencent Hy3 | $0.140 | $0.580 | 262K |
| DeepSeek Ai DeepSeek R1 Distill Llama 70B | $0.200 | $0.600 | 131K |
| Nvidia Llama 3.1 Nemotron 70B Instruct | $0.600 | $0.600 | 131K |
| OpenAI GPT Oss 120B Turbo | $0.150 | $0.600 | 131K |
| Qwen Qwen2.5 VL 32B Instruct | $0.200 | $0.600 | 128K |
| Qwen Qwen3 VL 30B A3b Instruct | $0.150 | $0.600 | 262K |
| Nousresearch Hermes 3 Llama 3.1 70B | $0.700 | $0.700 | 131K |
| Anthropic Claude 4 Opus | $16.50 | $82.50 | 200K |
Frequently asked questions
How much does DeepInfra cost per 1M tokens?
DeepInfra pricing starts at $0.020 per 1 million output tokens (Meta Llama Llama 3.2 3B Instruct) across 197 models, with its current flagship Qwen3.8 2.4T A95B at $6.00. Older premium models in the catalog list higher. Rates current as of August 31, 2026.
What is the cheapest DeepInfra model?
Meta Llama Llama 3.2 3B Instruct is the cheapest DeepInfra model on output tokens at $0.020 per 1M, while Mistralai Mistral Nemo Instruct 2407 is cheapest on input at $0.019 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest DeepInfra context window?
DeepSeek Ai DeepSeek V4 Flash has the largest context window in the DeepInfra catalog at 1.0M input tokens, with up to 1.0M output tokens per response.
Does DeepInfra support prompt caching?
Yes. GLM-4.7-Flash reads cached input at $0.010 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which DeepInfra models support reasoning or vision?
The DeepInfra catalog includes 50 reasoning models, 29 vision models, 61 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual DeepInfra bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator