DeepInfra Pricing
DeepInfra charges between $0.020 and $7.50 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; Qwen3.7 Max is the most expensive at $7.50/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Meta Llama Llama 3.2 3B Instruct
Cheapest output
$0.020/1M
Meta Llama Llama 3.2 3B Instruct
Longest context
1.0M
DeepSeek V4 Flash
Models priced
107
Avg $2.25/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 107 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Meta Llama Llama 3.2 3B Instruct | $0.020 | $0.020 | 131K |
| Meta Llama Meta Llama 3.1 8B Instruct Turbo | $0.020 | $0.030 | 131K |
| Mistralai Mistral Nemo Instruct 2407 | $0.020 | $0.040 | 131K |
| Meta Llama Llama 3.2 11B Vision Instruct | $0.049 | $0.049 | 131K |
| Meta Llama Meta Llama 3.1 8B Instruct | $0.030 | $0.050 | 131K |
| Sao10k L3 8B Lunaris V1 Turbo | $0.040 | $0.050 | 8K |
| Meta Llama Llama Guard 3 8B | $0.055 | $0.055 | 131K |
| Meta Llama Meta Llama 3 8B Instruct | $0.030 | $0.060 | 8K |
| Google Gemma 3 4B It | $0.040 | $0.080 | 131K |
| Mistralai Mistral Small 24B Instruct 2501 | $0.050 | $0.080 | 33K |
| Gryphe Mythomax L2 13B | $0.080 | $0.090 | 4K |
| Google Gemma 3 12B It | $0.050 | $0.100 | 131K |
| Qwen Qwen2.5 7B Instruct | $0.040 | $0.100 | 33K |
| GPT OSS 20B ReasoningTools | $0.030 | $0.140 | 131K |
| Microsoft Phi 4 | $0.070 | $0.140 | 16K |
| OpenAI GPT Oss 20B | $0.040 | $0.150 | 131K |
| Qwen3.5 9B VisionReasoningTools | $0.100 | $0.150 | 262K |
| Google Gemma 3 27B It | $0.090 | $0.160 | 131K |
| Nvidia Nvidia Nemotron Nano 9B | $0.040 | $0.160 | 131K |
| GPT OSS 120B ReasoningTools | $0.037 | $0.170 | 131K |
| DeepSeek V4 Flash ReasoningToolsCache | $0.090 | $0.180 | 1.0M |
| Meta Llama Llama Guard 4 12B | $0.180 | $0.180 | 164K |
| Mistralai Mistral Small 3.2 24B Instruct 2506 | $0.075 | $0.200 | 128K |
| Nemotron 3 Nano 30B A3B ReasoningTools | $0.050 | $0.200 | 262K |
| Qwen Qwen3 14B | $0.060 | $0.240 | 41K |
| DeepSeek Ai DeepSeek R1 Distill Qwen 32B | $0.270 | $0.270 | 131K |
| Meta Llama Meta Llama 3.1 70B Instruct Turbo | $0.100 | $0.280 | 131K |
| Qwen Qwen3 32B | $0.100 | $0.280 | 41K |
| Qwen3 32B ReasoningTools | $0.080 | $0.280 | 41K |
| Qwen Qwen3 30B A3b | $0.080 | $0.290 | 41K |
| Llama 4 Scout 17B VisionTools | $0.100 | $0.300 | 328K |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.080 | $0.300 | 328K |
| Nousresearch Hermes 3 Llama 3.1 70B | $0.300 | $0.300 | 131K |
| Llama 3.3 70B Turbo Tools | $0.100 | $0.320 | 131K |
| Gemma 4 26B A4B IT VisionReasoningTools | $0.070 | $0.340 | 262K |
| DeepSeek-V3.2 ReasoningToolsCache | $0.260 | $0.380 | 164K |
| Gemma 4 31B IT VisionReasoningTools | $0.130 | $0.380 | 262K |
| Meta Llama Llama 3.3 70B Instruct Turbo | $0.130 | $0.390 | 131K |
| Qwen Qwen2.5 72B Instruct | $0.120 | $0.390 | 33K |
| GLM-4.7-Flash ReasoningToolsCache | $0.060 | $0.400 | 203K |
| Google Gemini 2.0 Flash 001 | $0.100 | $0.400 | 1.0M |
| Llama 3.3 Nemotron Super 49B v1.5 ReasoningTools | $0.400 | $0.400 | 131K |
| Meta Llama Llama 3.3 70B Instruct | $0.230 | $0.400 | 131K |
| Meta Llama Meta Llama 3.1 70B Instruct | $0.400 | $0.400 | 131K |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.400 | $0.400 | 33K |
| Nvidia Llama 3.3 Nemotron Super 49B V1.5 | $0.100 | $0.400 | 131K |
| Qwen Qwq 32B | $0.150 | $0.400 | 131K |
| OpenAI GPT Oss 120B | $0.050 | $0.450 | 131K |
| Microsoft Wizardlm 2 8X22B | $0.480 | $0.480 | 66K |
| Qwen Qwen3 235B A22b | $0.180 | $0.540 | 41K |
| DeepSeek Ai DeepSeek R1 Distill Llama 70B | $0.200 | $0.600 | 131K |
| Meta Llama Llama 4 Maverick 17B 128e Instruct Fp8 | $0.150 | $0.600 | 1.0M |
| Nvidia Llama 3.1 Nemotron 70B Instruct | $0.600 | $0.600 | 131K |
| Qwen Qwen2.5 VL 32B Instruct | $0.200 | $0.600 | 128K |
| Qwen Qwen3 235B A22b Instruct 2507 | $0.090 | $0.600 | 262K |
| Sao10k L3.1 70B Euryale V2.2 | $0.650 | $0.750 | 131K |
| Sao10k L3.3 70B Euryale V2.3 | $0.650 | $0.750 | 131K |
| Llama 4 Maverick 17B FP8 Vision | $0.200 | $0.800 | 1.0M |
| Nemotron 3 Nano Omni 30B A3B Reasoning VisionReasoningTools | $0.200 | $0.800 | 262K |
| DeepSeek Ai DeepSeek V3 0324 | $0.250 | $0.880 | 164K |
| DeepSeek Ai DeepSeek V3 | $0.380 | $0.890 | 164K |
| Qwen3.6 35B A3B VisionReasoningTools | $0.150 | $0.950 | 262K |
| DeepSeek Ai DeepSeek V3.1 | $0.270 | $1.00 | 164K |
| DeepSeek Ai DeepSeek V3.1 Terminus | $0.270 | $1.00 | 164K |
| MiniMax-M2.7 ReasoningToolsCache | $0.250 | $1.00 | 197K |
| Nousresearch Hermes 3 Llama 3.1 405B | $1.00 | $1.00 | 131K |
| Qwen 3.5 35B A3B VisionReasoningToolsCache | $0.140 | $1.00 | 262K |
| Qwen3 Coder 480B A35B Instruct Turbo ToolsCache | $0.300 | $1.00 | 262K |
| Qwen3-Next 80B-A3B Instruct Tools | $0.090 | $1.10 | 262K |
| MiniMax M2.5 ReasoningToolsCache | $0.150 | $1.15 | 197K |
| MiniMax-M3 VisionReasoningToolsCache | $0.300 | $1.20 | 524K |
| Qwen Qwen3 Coder 480B A35b Instruct Turbo | $0.290 | $1.20 | 262K |
| Qwen Qwen3 Next 80B A3b Instruct | $0.140 | $1.40 | 262K |
| Qwen Qwen3 Next 80B A3b Thinking | $0.140 | $1.40 | 262K |
| Allenai Olmocr 7B 0725 Fp8 | $0.270 | $1.50 | 16K |
| Qwen Qwen3 Coder 480B A35b Instruct | $0.400 | $1.60 | 262K |
| Zai Org Glm 4.5 | $0.400 | $1.60 | 131K |
| GLM-4.7 ReasoningToolsCache | $0.400 | $1.75 | 203K |
| GLM-4.6 ReasoningToolsCache | $0.500 | $2.00 | 203K |
| Anthropic Claude 4 Opus | $16.50 | $82.50 | 200K |
Frequently asked questions
How much does DeepInfra cost per 1M tokens?
DeepInfra pricing ranges from $0.020 to $7.50 per 1 million output tokens across 107 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest DeepInfra model?
Meta Llama Llama 3.2 3B Instruct is the cheapest DeepInfra model on both axes — $0.020 per 1M input tokens and $0.020 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest DeepInfra context window?
DeepSeek V4 Flash has the largest context window in the DeepInfra catalog at 1.0M input tokens, with up to 16K output tokens per response.
Does DeepInfra support prompt caching?
Yes. GLM-4.7-Flash reads cached input at $0.010 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which DeepInfra models support reasoning or vision?
The DeepInfra catalog includes 33 reasoning models, 17 vision models, 39 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual DeepInfra bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator