Azure OpenAI Pricing
Azure OpenAI charges from $0.040 per 1 million output tokens, depending on the model. Ministral 3B is the cheapest at $0.040/1M output and $0.040/1M input; its current flagship GPT-5.6 Sol costs $30.00/1M output. The largest context window is 1.1M tokens (Eu GPT 5.4). Prices are USD, current as of August 11, 2026.
Azure OpenAI carries the OpenAI model lineup under Azure billing, with regional availability and quota managed per deployment. Standard (pay-as-you-go) token rates are listed below; provisioned throughput units are priced separately.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.040/1M
Ministral 3B
Cheapest output
$0.040/1M
Ministral 3B
Longest context
1.1M
Eu GPT 5.4
Models priced
150
Avg $17.36/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 150 priced models, cheapest first. The batch column is the asynchronous batch-API rate, 50% below the standard output rate on the 14 models that publish one.
| Model | Input / 1M | Output / 1M | Batch out / 1M | Context |
|---|---|---|---|---|
| Model Router VisionTools | $0.140 | — | — | 200K |
| Ministral 3B Tools | $0.040 | $0.040 | — | 128K |
| Mistral Small 3.1 VisionTools | $0.100 | $0.300 | — | 128K |
| Phi-4-mini Tools | $0.075 | $0.300 | — | 128K |
| Phi-4-mini-reasoning ReasoningTools | $0.075 | $0.300 | — | 128K |
| Phi-4-multimodal Vision | $0.080 | $0.320 | — | 128K |
| GPT-4.1 nano VisionToolsCache | $0.100 | $0.400 | $0.200−50% | 1.0M |
| GPT-5 Nano VisionReasoningToolsCache | $0.050 | $0.400 | — | 400K |
| Eu GPT 5 Nano 2025 08.07 | $0.055 | $0.440 | — | 272K |
| Us GPT 4.1 Nano 2025 04.14 | $0.110 | $0.440 | $0.220−50% | 1.0M |
| Us GPT 5 Nano 2025 08.07 | $0.055 | $0.440 | — | 272K |
| Grok 4.1 Fast (Non-Reasoning) VisionToolsCache | $0.200 | $0.500 | — | 128K |
| Grok 4.1 Fast (Reasoning) VisionReasoningToolsCache | $0.200 | $0.500 | — | 128K |
| Phi-4 | $0.125 | $0.500 | — | 128K |
| Phi-4-reasoning Reasoning | $0.125 | $0.500 | — | 32K |
| Phi-4-reasoning-plus Reasoning | $0.125 | $0.500 | — | 32K |
| DeepSeek-V4-Flash Reasoning | $0.190 | $0.510 | — | 1.0M |
| Global Standard GPT 4o Mini | $0.150 | $0.600 | — | 128K |
| GPT-4o mini VisionToolsCache | $0.150 | $0.600 | — | 128K |
| Eu GPT 4o Mini 2024 07.18 | $0.165 | $0.660 | — | 128K |
| GPT 4o Mini 2024 07.18 | $0.165 | $0.660 | — | 128K |
| Us GPT 4o Mini 2024 07.18 | $0.165 | $0.660 | — | 128K |
| Llama-3.3-70B-Instruct Tools | $0.710 | $0.710 | — | 128K |
| Llama 4 Scout 17B 16E Instruct VisionTools | $0.200 | $0.780 | — | 128K |
| Codestral 25.01 Tools | $0.300 | $0.900 | — | 256K |
| Llama 4 Maverick 17B 128E Instruct FP8 VisionTools | $0.250 | $1.00 | — | 1.0M |
| GPT-5.4 Nano VisionReasoningToolsCache | $0.200 | $1.25 | — | 400K |
| Eu GPT 5.6 Luna | $0.220 | $1.32 | — | 1.1M |
| Us GPT 5.6 Luna | $0.220 | $1.32 | — | 1.1M |
| GPT 3.5 Turbo | $0.500 | $1.50 | — | 4K |
| GPT-3.5 Turbo 0125 | $0.500 | $1.50 | — | 16K |
| GPT-4.1 mini VisionToolsCache | $0.400 | $1.60 | $0.800−50% | 1.0M |
| DeepSeek-V3.2 ReasoningTools | $0.580 | $1.68 | — | 128K |
| DeepSeek-V3.2-Speciale Reasoning | $0.580 | $1.68 | — | 128K |
| Us GPT 4.1 Mini 2025 04.14 | $0.440 | $1.76 | $0.880−50% | 1.0M |
| GPT-3.5 Turbo 1106 | $1.00 | $2.00 | — | 16K |
| GPT-3.5 Turbo Instruct | $1.50 | $2.00 | — | 4K |
| GPT-5 Mini VisionReasoningToolsCache | $0.250 | $2.00 | — | 400K |
| GPT-5.1 Codex Mini VisionReasoningToolsCache | $0.250 | $2.00 | — | 400K |
| Mistral Medium 3 VisionTools | $0.400 | $2.00 | — | 128K |
| Eu GPT 5 Mini 2025 08.07 | $0.275 | $2.20 | — | 272K |
| Us GPT 5 Mini 2025 08.07 | $0.275 | $2.20 | — | 272K |
| GPT Audio Mini 2025 10.06 | $0.600 | $2.40 | — | 128K |
| Kimi K2.5 VisionReasoningTools | $0.600 | $3.00 | — | 262K |
| DeepSeek-V4-Pro Reasoning | $1.74 | $3.48 | — | 1.0M |
| GPT 35 Turbo 16K | $3.00 | $4.00 | — | 16K |
| GPT 35 Turbo 16K 0613 | $3.00 | $4.00 | — | 16K |
| Kimi K2.6 VisionReasoningTools | $0.950 | $4.00 | — | 262K |
| Kimi K2.7 Code VisionReasoningToolsCache | $0.950 | $4.00 | — | 262K |
| o1 Mini 2024 09.12 | $1.10 | $4.40 | — | 128K |
| o3-mini ReasoningToolsCache | $1.10 | $4.40 | — | 200K |
| o4-mini VisionReasoningToolsCache | $1.10 | $4.40 | — | 200K |
| GPT-5.4 Mini VisionReasoningToolsCache | $0.750 | $4.50 | — | 400K |
| Eu o1 Mini 2024 09.12 | $1.21 | $4.84 | $2.42−50% | 128K |
| Eu o3 Mini 2025 01.31 | $1.21 | $4.84 | $2.42−50% | 200K |
| o1 Mini | $1.21 | $4.84 | — | 128K |
| Us o1 Mini 2024 09.12 | $1.21 | $4.84 | $2.42−50% | 128K |
| Us o3 Mini 2025 01.31 | $1.21 | $4.84 | $2.42−50% | 200K |
| Us o4 Mini 2025 04.16 | $1.21 | $4.84 | — | 200K |
| Claude Haiku 4.5 VisionReasoningToolsCache | $1.00 | $5.00 | — | 200K |
| DeepSeek-R1 Reasoning | $1.35 | $5.40 | — | 164K |
| Codex Mini ReasoningToolsCache | $1.50 | $6.00 | — | 200K |
| GPT-5.6 Luna VisionReasoningToolsCache | $1.00 | $6.00 | — | 1.1M |
| Grok 4.20 (Non-Reasoning) Tools | $2.00 | $6.00 | — | 262K |
| Grok 4.20 (Reasoning) ReasoningTools | $2.00 | $6.00 | — | 262K |
| GPT-4.1 VisionToolsCache | $2.00 | $8.00 | $4.00−50% | 1.0M |
| o3 VisionReasoningToolsCache | $2.00 | $8.00 | — | 200K |
| Us GPT 4.1 2025 04.14 | $2.20 | $8.80 | $4.40−50% | 1.0M |
| Us o3 2025 04.16 | $2.20 | $8.80 | — | 200K |
| Claude Sonnet 5 VisionReasoningToolsCache | $2.00 | $10.00 | — | 1.0M |
| Command A ReasoningTools | $2.50 | $10.00 | — | 131K |
| Global GPT 4o 2024 08.06 | $2.50 | $10.00 | — | 128K |
| Global GPT 5.1 | $1.25 | $10.00 | — | 272K |
| Global GPT 5.1 Chat | $1.25 | $10.00 | — | 128K |
| Global Standard GPT 4o 2024 08.06 | $2.50 | $10.00 | — | 128K |
| GPT 4o Audio Preview 2024 12.17 | $2.50 | $10.00 | — | 128K |
| GPT 4o Mini Audio Preview 2024 12.17 | $2.50 | $10.00 | — | 128K |
| GPT 5 Chat | $1.25 | $10.00 | — | 128K |
| GPT 5 Chat Latest | $1.25 | $10.00 | — | 128K |
| GPT-5.4 Pro VisionReasoningTools | $30.00 | $180.00 | — | 1.1M |
A dash in the batch column means our catalog carries no published batch rate for that model — not that the model has no batch tier. Batch coverage in the upstream pricing data is uneven.
Frequently asked questions
How much does Azure OpenAI cost per 1M tokens?
Azure OpenAI pricing starts at $0.040 per 1 million output tokens (Ministral 3B) across 150 models, with its current flagship GPT-5.6 Sol at $30.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Azure OpenAI model?
Ministral 3B is the cheapest Azure OpenAI model on both axes — $0.040 per 1M input tokens and $0.040 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Azure OpenAI context window?
Eu GPT 5.4 has the largest context window in the Azure OpenAI catalog at 1.1M input tokens, with up to 128K output tokens per response.
Does Azure OpenAI support prompt caching?
Yes. GPT-5 Nano reads cached input at $0.010 per 1M tokens, against $0.050 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Does Azure OpenAI offer batch pricing?
Yes — 14 Azure OpenAI models in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Us GPT 4.1 2025 04.14 bills $4.40 per 1M output in batch against $8.80 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.
Which Azure OpenAI models support reasoning or vision?
The Azure OpenAI catalog includes 50 reasoning models, 53 vision models, 62 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Azure OpenAI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator