Azure AI Pricing
Azure AI charges between — and $9710.00 per 1 million output tokens, depending on the model. Model Router is the cheapest at —/1M output and $0.140/1M input; Jais 30B Chat is the most expensive at $9710.00/1M output. The largest context window is 10.0M tokens (Llama 4 Scout 17B 16e Instruct). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.040/1M
Ministral 3B
Cheapest output
—/1M
Model Router
Longest context
10.0M
Llama 4 Scout 17B 16e Instruct
Models priced
71
Avg $136.48/1M output
Pricing by model
All prices in USD per 1 million tokens. All 71 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Model Router | $0.140 | — | 0 |
| Ministral 3B | $0.040 | $0.040 | 128K |
| Mistral Nemo | $0.150 | $0.150 | 131K |
| Mistral Small 2503 | $0.100 | $0.300 | 128K |
| Phi 4 Mini Instruct | $0.075 | $0.300 | 131K |
| Phi 4 Mini Reasoning | $0.080 | $0.320 | 131K |
| Phi 4 Multimodal Instruct | $0.080 | $0.320 | 131K |
| Llama 4 Maverick 17B 128e Instruct Fp8 | $1.41 | $0.350 | 1.0M |
| Llama 3.2 11B Vision Instruct | $0.370 | $0.370 | 128K |
| Meta Llama 3 70B Instruct | $1.10 | $0.370 | 8K |
| Grok 4 Fast Non Reasoning | $0.200 | $0.500 | 131K |
| Grok 4 Fast Reasoning | $0.200 | $0.500 | 131K |
| Grok 4.1 Fast Non Reasoning | $0.200 | $0.500 | 131K |
| Grok 4.1 Fast Reasoning | $0.200 | $0.500 | 131K |
| Phi 4 | $0.125 | $0.500 | 16K |
| Phi 4 Reasoning | $0.125 | $0.500 | 33K |
| DeepSeek V4 Flash | $0.190 | $0.510 | 1.0M |
| Phi 3 Mini 128K Instruct | $0.130 | $0.520 | 128K |
| Phi 3 Mini 4K Instruct | $0.130 | $0.520 | 4K |
| Phi 3.5 Mini Instruct | $0.130 | $0.520 | 128K |
| Phi 3.5 Vision Instruct | $0.130 | $0.520 | 128K |
| GPT Oss 120B | $0.150 | $0.600 | 131K |
| Phi 3 Small 128K Instruct | $0.150 | $0.600 | 128K |
| Phi 3 Small 8K Instruct | $0.150 | $0.600 | 8K |
| Meta Llama 3.1 8B Instruct | $0.300 | $0.610 | 128K |
| Phi 3.5 Moe Instruct | $0.160 | $0.640 | 128K |
| Phi 3 Medium 128K Instruct | $0.170 | $0.680 | 128K |
| Phi 3 Medium 4K Instruct | $0.170 | $0.680 | 4K |
| Jamba Instruct | $0.500 | $0.700 | 70K |
| Llama 3.3 70B Instruct | $0.710 | $0.710 | 128K |
| Llama 4 Scout 17B 16e Instruct | $0.200 | $0.780 | 10.0M |
| GPT 5.4 Nano | $0.200 | $1.25 | 400K |
| Global Grok 3 Mini | $0.250 | $1.27 | 131K |
| Grok 3 Mini | $0.250 | $1.27 | 131K |
| Grok Code Fast 1 | $0.200 | $1.50 | 131K |
| Mistral Large 3 | $0.500 | $1.50 | 256K |
| DeepSeek V3.2 | $0.580 | $1.68 | 164K |
| DeepSeek V3.2 Speciale | $0.580 | $1.68 | 164K |
| Mistral Medium 2505 | $0.400 | $2.00 | 131K |
| Llama 3.2 90B Vision Instruct | $2.04 | $2.04 | 128K |
| Kimi K2.5 | $0.600 | $3.00 | 262K |
| Mistral Small | $1.00 | $3.00 | 32K |
| DeepSeek V4 Pro | $1.74 | $3.48 | 1.0M |
| Meta Llama 3.1 70B Instruct | $2.68 | $3.54 | 128K |
| Kimi K2.6 | $0.950 | $4.00 | 262K |
| GPT 5.4 Mini | $0.750 | $4.50 | 400K |
| DeepSeek | $1.14 | $4.56 | 128K |
| DeepSeek V3 0324 | $1.14 | $4.56 | 128K |
| DeepSeek V3.1 | $1.23 | $4.94 | 131K |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| DeepSeek R1 | $1.35 | $5.40 | 128K |
| Mai Ds R1 | $1.35 | $5.40 | 128K |
| Mistral Large 2407 | $2.00 | $6.00 | 128K |
| Mistral Large Latest | $2.00 | $6.00 | 128K |
| Claude Sonnet 5 | $2.00 | $10.00 | 1.0M |
| Mistral Large | $4.00 | $12.00 | 32K |
| Claude Sonnet 4.5 | $3.00 | $15.00 | 200K |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1.0M |
| Global Grok 3 | $3.00 | $15.00 | 131K |
| GPT 5.4 | $2.50 | $15.00 | 1.1M |
| Grok 3 | $3.00 | $15.00 | 131K |
| Grok 4 | $3.00 | $15.00 | 131K |
| Meta Llama 3.1 405B Instruct | $5.33 | $16.00 | 128K |
| Claude Opus 4.5 | $5.00 | $25.00 | 200K |
| Claude Opus 4.6 | $5.00 | $25.00 | 200K |
| Claude Opus 4.7 | $5.00 | $25.00 | 200K |
| Claude Opus 4.8 | $5.00 | $25.00 | 200K |
| GPT 5.5 | $5.00 | $30.00 | 1.1M |
| Claude Fable 5 | $10.00 | $50.00 | 1.0M |
| Claude Opus 4.1 | $15.00 | $75.00 | 200K |
| Jais 30B Chat | $3200.00 | $9710.00 | 8K |
Frequently asked questions
How much does Azure AI cost per 1M tokens?
Azure AI pricing ranges from — to $9710.00 per 1 million output tokens across 71 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Azure AI model?
Model Router is the cheapest Azure AI model on output tokens at — per 1M, while Ministral 3B is cheapest on input at $0.040 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Azure AI context window?
Llama 4 Scout 17B 16e Instruct has the largest context window in the Azure AI catalog at 10.0M input tokens, with up to 16K output tokens per response.
How do I calculate my actual Azure AI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator