API pricing

Azure AI Pricing

Azure AI charges from $0.040 per 1 million output tokens, depending on the model. Ministral 3B is the cheapest at $0.040/1M output and $0.040/1M input; its current flagship Jais 30B Chat costs $9710.00/1M output. The largest context window is 10.0M tokens (Llama 4 Scout 17B 16e Instruct). Prices are USD, current as of August 11, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.040/1M

Ministral 3B

Cheapest output

$0.040/1M

Ministral 3B

Longest context

10.0M

Llama 4 Scout 17B 16e Instruct

Models priced

72

Avg $135.01/1M output

Pricing by model

All prices in USD per 1 million tokens. All 72 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Model Router$0.1400
Ministral 3B$0.040$0.040128K
Mistral Nemo$0.150$0.150131K
Mistral Small 2503$0.100$0.300128K
Phi 4 Mini Instruct$0.075$0.300131K
Phi 4 Mini Reasoning$0.080$0.320131K
Phi 4 Multimodal Instruct$0.080$0.320131K
Llama 4 Maverick 17B 128e Instruct Fp8$1.41$0.3501.0M
Llama 3.2 11B Vision Instruct$0.370$0.370128K
Meta Llama 3 70B Instruct$1.10$0.3708K
Grok 4 Fast Non Reasoning$0.200$0.500131K
Grok 4 Fast Reasoning$0.200$0.500131K
Grok 4.1 Fast Non Reasoning$0.200$0.500131K
Grok 4.1 Fast Reasoning$0.200$0.500131K
Phi 4$0.125$0.50016K
Phi 4 Reasoning$0.125$0.50033K
DeepSeek V4 Flash$0.190$0.5101.0M
Phi 3 Mini 128K Instruct$0.130$0.520128K
Phi 3 Mini 4K Instruct$0.130$0.5204K
Phi 3.5 Mini Instruct$0.130$0.520128K
Phi 3.5 Vision Instruct$0.130$0.520128K
GPT Oss 120B$0.150$0.600131K
Phi 3 Small 128K Instruct$0.150$0.600128K
Phi 3 Small 8K Instruct$0.150$0.6008K
Meta Llama 3.1 8B Instruct$0.300$0.610128K
Phi 3.5 Moe Instruct$0.160$0.640128K
Phi 3 Medium 128K Instruct$0.170$0.680128K
Phi 3 Medium 4K Instruct$0.170$0.6804K
Jamba Instruct$0.500$0.70070K
Llama 3.3 70B Instruct$0.710$0.710128K
Llama 4 Scout 17B 16e Instruct$0.200$0.78010.0M
GPT 5.4 Nano$0.200$1.25272K
Global Grok 3 Mini$0.250$1.27131K
Grok 3 Mini$0.250$1.27131K
Grok Code Fast 1$0.200$1.50131K
Mistral Large 3$0.500$1.50256K
DeepSeek V3.2$0.580$1.68164K
DeepSeek V3.2 Speciale$0.580$1.68164K
Mistral Medium 2505$0.400$2.00131K
Llama 3.2 90B Vision Instruct$2.04$2.04128K
Kimi K2.5$0.600$3.00262K
Mistral Small$1.00$3.0032K
DeepSeek V4 Pro$1.74$3.481.0M
Meta Llama 3.1 70B Instruct$2.68$3.54128K
Kimi K2.6$0.950$4.00262K
GPT 5.4 Mini$0.750$4.50272K
DeepSeek$1.14$4.56128K
DeepSeek V3 0324$1.14$4.56128K
DeepSeek V3.1$1.23$4.94131K
Claude Haiku 4.5$1.00$5.00200K
DeepSeek R1$1.35$5.40128K
Mai Ds R1$1.35$5.40128K
Mistral Large 2407$2.00$6.00128K
Mistral Large Latest$2.00$6.00128K
Claude Sonnet 5$2.00$10.001.0M
Mistral Large$4.00$12.0032K
Claude Sonnet 4.5$3.00$15.00200K
Claude Sonnet 4.6$3.00$15.001.0M
Global Grok 3$3.00$15.00131K
GPT 5.4$2.50$15.001.1M
Grok 3$3.00$15.00131K
Grok 4$3.00$15.00131K
Meta Llama 3.1 405B Instruct$5.33$16.00128K
Claude Opus 4.5$5.00$25.00200K
Claude Opus 4.6$5.00$25.001.0M
Claude Opus 4.7$5.00$25.001.0M
Claude Opus 4.8$5.00$25.001.0M
Claude Opus 5$5.00$25.001.0M
GPT 5.5$5.00$30.001.1M
Claude Fable 5$10.00$50.001.0M
Claude Opus 4.1$15.00$75.00200K
Jais 30B Chat$3200.00$9710.008K

Frequently asked questions

How much does Azure AI cost per 1M tokens?

Azure AI pricing starts at $0.040 per 1 million output tokens (Ministral 3B) across 72 models, with its current flagship Jais 30B Chat at $9710.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.

What is the cheapest Azure AI model?

Ministral 3B is the cheapest Azure AI model on both axes — $0.040 per 1M input tokens and $0.040 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Azure AI context window?

Llama 4 Scout 17B 16e Instruct has the largest context window in the Azure AI catalog at 10.0M input tokens, with up to 16K output tokens per response.

How do I calculate my actual Azure AI bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator