API pricing

Vercel AI Gateway Pricing

Vercel AI Gateway charges from $0.040 per 1 million output tokens, depending on the model. Mistral Ministral 3B is the cheapest at $0.040/1M output and $0.040/1M input; its current flagship Anthropic Claude 3 Opus costs $75.00/1M output. The largest context window is 1.0M tokens (Google Gemini 2.0 Flash). Prices are USD, current as of August 11, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.020/1M

Amazon Titan Embed Text

Cheapest output

$0.040/1M

Mistral Ministral 3B

Longest context

1.0M

Google Gemini 2.0 Flash

Models priced

90

Avg $8.63/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 90 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Amazon Titan Embed Text$0.0200
Cohere Embed V4.0$0.1200
Mistral Codestral Embed$0.1500
Mistral Mistral Embed$0.1000
Mistral Ministral 3B$0.040$0.040128K
Meta Llama 3 8B$0.050$0.0808K
Meta Llama 3.1 8B$0.050$0.080131K
Meta Llama 3.2 1B$0.100$0.100128K
Mistral Ministral 8B$0.100$0.100128K
Amazon Nova Micro$0.035$0.140128K
Meta Llama 3.2 3B$0.150$0.150128K
Mistral Pixtral 12B$0.150$0.150128K
Meta Llama 3.2 11B$0.160$0.160128K
Google Gemma 2 9B$0.200$0.2008K
Alibaba Qwen 3 14B$0.080$0.24041K
Amazon Nova Lite$0.060$0.240300K
Mistral Devstral Small$0.070$0.280128K
Alibaba Qwen 3 30B$0.100$0.30041K
Alibaba Qwen 3 32B$0.100$0.30041K
Google Gemini 2.0 Flash Lite$0.075$0.3001.0M
Meta Llama 4 Scout$0.100$0.300131K
Mistral Mistral Small$0.100$0.30032K
OpenAI GPT 4.1 Nano$0.100$0.4001.0M
Xai Grok 3 Mini$0.300$0.500131K
Alibaba Qwen 3 235B$0.200$0.60041K
Cohere Command R$0.150$0.600128K
Google Gemini 2.0 Flash$0.150$0.6001.0M
Meta Llama 4 Maverick$0.200$0.600131K
OpenAI GPT 4o Mini$0.150$0.600128K
Meta Llama 3.1 70B$0.720$0.720128K
Meta Llama 3.2 90B$0.720$0.720128K
Meta Llama 3.3 70B$0.720$0.720128K
Meta Llama 3 70B$0.590$0.7908K
Mistral Mistral Saba 24B$0.790$0.79033K
DeepSeek DeepSeek$0.900$0.900128K
Mistral Codestral$0.300$0.900256K
DeepSeek DeepSeek R1 Distill Llama 70B$0.750$0.990131K
Inception Mercury Coder Small$0.250$1.0032K
Perplexity Sonar$1.00$1.00127K
Zai Glm 4.5 Air$0.200$1.10128K
Mistral Mixtral 8X22B Instruct$1.20$1.2066K
Morph Morph V3 Fast$0.800$1.2033K
Anthropic Claude 3 Haiku$0.250$1.25200K
Mistral Magistral Small$0.500$1.50128K
OpenAI GPT 3.5 Turbo$0.500$1.5016K
Alibaba Qwen3 Coder$0.400$1.60262K
OpenAI GPT 4.1 Mini$0.400$1.601.0M
Zai Glm 4.6$0.450$1.80200K
Morph Morph V3 Large$0.900$1.9033K
OpenAI GPT 3.5 Turbo Instruct$1.50$2.008K
DeepSeek DeepSeek R1$0.550$2.19128K
Moonshotai Kimi K2$0.550$2.20131K
Zai Glm 4.5$0.600$2.20131K
Google Gemini 2.5 Flash$0.300$2.501.0M
Amazon Nova Pro$0.800$3.20300K
Anthropic Claude 3.5 Haiku$0.800$4.00200K
Xai Grok 3 Mini Fast$0.600$4.00131K
OpenAI o3 Mini$1.10$4.40200K
OpenAI o4 Mini$1.10$4.40200K
Anthropic Claude Haiku 4.5$1.00$5.00200K
Mistral Magistral Medium$2.00$5.00128K
Perplexity Sonar Reasoning$1.00$5.00127K
Mistral Mistral Large$2.00$6.0032K
Mistral Pixtral Large$2.00$6.00128K
OpenAI GPT 4.1$2.00$8.001.0M
OpenAI o3$2.00$8.00200K
Perplexity Sonar Reasoning Pro$2.00$8.00127K
Cohere Command A$2.50$10.00256K
Cohere Command R Plus$2.50$10.00128K
Google Gemini 2.5 Pro$2.50$10.001.0M
OpenAI GPT 4o$2.50$10.00128K
Xai Grok 2$2.00$10.00131K
Xai Grok 2 Vision$2.00$10.0033K
Anthropic Claude 3.5 Sonnet$3.00$15.00200K
Anthropic Claude 3.7 Sonnet$3.00$15.00200K
Anthropic Claude 4 Sonnet$3.00$15.00200K
Anthropic Claude Sonnet 4.5$3.00$15.001.0M
Perplexity Sonar Pro$3.00$15.00200K
Vercel V0 1.0 Md$3.00$15.00128K
Anthropic Claude Opus 4.1$15.00$75.00200K

Frequently asked questions

How much does Vercel AI Gateway cost per 1M tokens?

Vercel AI Gateway pricing starts at $0.040 per 1 million output tokens (Mistral Ministral 3B) across 90 models, with its current flagship Anthropic Claude 3 Opus at $75.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.

What is the cheapest Vercel AI Gateway model?

Mistral Ministral 3B is the cheapest Vercel AI Gateway model on output tokens at $0.040 per 1M, while Amazon Titan Embed Text is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Vercel AI Gateway context window?

Google Gemini 2.0 Flash has the largest context window in the Vercel AI Gateway catalog at 1.0M input tokens, with up to 8K output tokens per response.

How do I calculate my actual Vercel AI Gateway bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator