Azure OpenAI Pricing
Azure OpenAI charges between — and $50.00 per 1 million output tokens, depending on the model. Model Router is the cheapest at —/1M output and $0.140/1M input; Claude Fable 5 is the most expensive at $50.00/1M output. The largest context window is 2.0M tokens (Grok 4 Fast (Reasoning)). Prices are USD, current as of July 20, 2026.
Azure OpenAI carries the OpenAI model lineup under Azure billing, with regional availability and quota managed per deployment. Standard (pay-as-you-go) token rates are listed below; provisioned throughput units are priced separately.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.040/1M
Ministral 3B
Cheapest output
—/1M
Model Router
Longest context
2.0M
Grok 4 Fast (Reasoning)
Models priced
173
Avg $15.88/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 173 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Model Router VisionTools | $0.140 | — | 200K |
| Ministral 3B Tools | $0.040 | $0.040 | 128K |
| Mistral Nemo Tools | $0.150 | $0.150 | 128K |
| Mistral Small 3.1 VisionTools | $0.100 | $0.300 | 128K |
| Phi-4-mini Tools | $0.075 | $0.300 | 128K |
| Phi-4-mini-reasoning ReasoningTools | $0.075 | $0.300 | 128K |
| Phi-4-multimodal Vision | $0.080 | $0.320 | 128K |
| Llama-3.2-11B-Vision-Instruct VisionTools | $0.370 | $0.370 | 128K |
| GPT-4.1 nano VisionToolsCache | $0.100 | $0.400 | 1.0M |
| GPT-5 Nano VisionReasoningToolsCache | $0.050 | $0.400 | 400K |
| Eu GPT 5 Nano 2025 08.07 | $0.055 | $0.440 | 272K |
| Us GPT 4.1 Nano 2025 04.14 | $0.110 | $0.440 | 1.0M |
| Us GPT 5 Nano 2025 08.07 | $0.055 | $0.440 | 272K |
| Grok 4 Fast (Reasoning) VisionReasoningToolsCache | $0.200 | $0.500 | 2.0M |
| Grok 4.1 Fast (Non-Reasoning) VisionToolsCache | $0.200 | $0.500 | 128K |
| Grok 4.1 Fast (Reasoning) VisionReasoningToolsCache | $0.200 | $0.500 | 128K |
| Phi-4 | $0.125 | $0.500 | 128K |
| Phi-4-reasoning Reasoning | $0.125 | $0.500 | 32K |
| Phi-4-reasoning-plus Reasoning | $0.125 | $0.500 | 32K |
| DeepSeek-V4-Flash Reasoning | $0.190 | $0.510 | 1.0M |
| Phi-3-mini-instruct (128k) | $0.130 | $0.520 | 128K |
| Phi-3-mini-instruct (4k) | $0.130 | $0.520 | 4K |
| Phi-3.5-mini-instruct | $0.130 | $0.520 | 128K |
| Command R Tools | $0.150 | $0.600 | 128K |
| Global Standard GPT 4o Mini | $0.150 | $0.600 | 128K |
| GPT-4o mini VisionToolsCache | $0.150 | $0.600 | 128K |
| Phi-3-small-instruct (128k) | $0.150 | $0.600 | 128K |
| Phi-3-small-instruct (8k) | $0.150 | $0.600 | 8K |
| Meta-Llama-3-8B-Instruct | $0.300 | $0.610 | 8K |
| Meta-Llama-3.1-8B-Instruct Tools | $0.300 | $0.610 | 128K |
| Phi-3.5-MoE-instruct | $0.160 | $0.640 | 128K |
| Eu GPT 4o Mini 2024 07.18 | $0.165 | $0.660 | 128K |
| GPT 4o Mini 2024 07.18 | $0.165 | $0.660 | 128K |
| Us GPT 4o Mini 2024 07.18 | $0.165 | $0.660 | 128K |
| Phi-3-medium-instruct (128k) | $0.170 | $0.680 | 128K |
| Phi-3-medium-instruct (4k) | $0.170 | $0.680 | 4K |
| Llama-3.3-70B-Instruct Tools | $0.710 | $0.710 | 128K |
| Llama 4 Scout 17B 16E Instruct VisionTools | $0.200 | $0.780 | 128K |
| Codestral 25.01 Tools | $0.300 | $0.900 | 256K |
| Llama 4 Maverick 17B 128E Instruct FP8 VisionTools | $0.250 | $1.00 | 1.0M |
| GPT-5.4 Nano VisionReasoningToolsCache | $0.200 | $1.25 | 400K |
| GPT 3.5 Turbo | $0.500 | $1.50 | 4K |
| GPT-3.5 Turbo 0125 | $0.500 | $1.50 | 16K |
| GPT-4.1 mini VisionToolsCache | $0.400 | $1.60 | 1.0M |
| DeepSeek-V3.1 ReasoningTools | $0.560 | $1.68 | 131K |
| DeepSeek-V3.2 ReasoningTools | $0.580 | $1.68 | 128K |
| DeepSeek-V3.2-Speciale Reasoning | $0.580 | $1.68 | 128K |
| Us GPT 4.1 Mini 2025 04.14 | $0.440 | $1.76 | 1.0M |
| GPT-3.5 Turbo 0301 | $1.50 | $2.00 | 4K |
| GPT-3.5 Turbo 1106 | $1.00 | $2.00 | 16K |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 | 4K |
| GPT-5 Mini VisionReasoningToolsCache | $0.250 | $2.00 | 400K |
| GPT-5.1 Codex Mini VisionReasoningToolsCache | $0.250 | $2.00 | 400K |
| Mistral Medium 3 VisionTools | $0.400 | $2.00 | 128K |
| Llama-3.2-90B-Vision-Instruct VisionTools | $2.04 | $2.04 | 128K |
| Eu GPT 5 Mini 2025 08.07 | $0.275 | $2.20 | 272K |
| Us GPT 5 Mini 2025 08.07 | $0.275 | $2.20 | 272K |
| GPT Audio Mini 2025 10.06 | $0.600 | $2.40 | 128K |
| Kimi K2 Thinking ReasoningToolsCache | $0.600 | $2.50 | 262K |
| Kimi K2.5 VisionReasoningTools | $0.600 | $3.00 | 262K |
| DeepSeek-V4-Pro Reasoning | $1.74 | $3.48 | 1.0M |
| Meta-Llama-3-70B-Instruct | $2.68 | $3.54 | 8K |
| Meta-Llama-3.1-70B-Instruct Tools | $2.68 | $3.54 | 128K |
| GPT 35 Turbo 16K | $3.00 | $4.00 | 16K |
| GPT 35 Turbo 16K 0613 | $3.00 | $4.00 | 16K |
| GPT-3.5 Turbo 0613 | $3.00 | $4.00 | 16K |
| Kimi K2.6 VisionReasoningTools | $0.950 | $4.00 | 262K |
| o1-mini ReasoningToolsCache | $1.10 | $4.40 | 128K |
| o3-mini ReasoningToolsCache | $1.10 | $4.40 | 200K |
| o4-mini VisionReasoningToolsCache | $1.10 | $4.40 | 200K |
| GPT-5.4 Mini VisionReasoningToolsCache | $0.750 | $4.50 | 400K |
| DeepSeek-V3-0324 Tools | $1.14 | $4.56 | 131K |
| Eu o1 Mini 2024 09.12 | $1.21 | $4.84 | 128K |
| Eu o3 Mini 2025 01.31 | $1.21 | $4.84 | 200K |
| Us o1 Mini 2024 09.12 | $1.21 | $4.84 | 128K |
| Us o3 Mini 2025 01.31 | $1.21 | $4.84 | 200K |
| Us o4 Mini 2025 04.16 | $1.21 | $4.84 | 200K |
| Claude Haiku 4.5 VisionReasoningToolsCache | $1.00 | $5.00 | 200K |
| DeepSeek-R1 Reasoning | $1.35 | $5.40 | 164K |
| GPT-5.4 Pro VisionReasoningTools | $30.00 | $180.00 | 1.1M |
Frequently asked questions
How much does Azure OpenAI cost per 1M tokens?
Azure OpenAI pricing ranges from — to $50.00 per 1 million output tokens across 173 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Azure OpenAI model?
Model Router is the cheapest Azure OpenAI model on output tokens at — per 1M, while Ministral 3B is cheapest on input at $0.040 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Azure OpenAI context window?
Grok 4 Fast (Reasoning) has the largest context window in the Azure OpenAI catalog at 2.0M input tokens, with up to 30K output tokens per response.
Does Azure OpenAI support prompt caching?
Yes. GPT-5 Nano reads cached input at $0.010 per 1M tokens, against $0.050 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Azure OpenAI models support reasoning or vision?
The Azure OpenAI catalog includes 57 reasoning models, 58 vision models, 80 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Azure OpenAI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator