Vercel AI Gateway Pricing
Vercel AI Gateway charges from $0.040 per 1 million output tokens, depending on the model. Mistral Ministral 3B is the cheapest at $0.040/1M output and $0.040/1M input; its current flagship Anthropic Claude 3 Opus costs $75.00/1M output. The largest context window is 1.0M tokens (Google Gemini 2.0 Flash). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Amazon Titan Embed Text
Cheapest output
$0.040/1M
Mistral Ministral 3B
Longest context
1.0M
Google Gemini 2.0 Flash
Models priced
90
Avg $8.63/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 90 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Amazon Titan Embed Text | $0.020 | — | 0 |
| Cohere Embed V4.0 | $0.120 | — | 0 |
| Mistral Codestral Embed | $0.150 | — | 0 |
| Mistral Mistral Embed | $0.100 | — | 0 |
| Mistral Ministral 3B | $0.040 | $0.040 | 128K |
| Meta Llama 3 8B | $0.050 | $0.080 | 8K |
| Meta Llama 3.1 8B | $0.050 | $0.080 | 131K |
| Meta Llama 3.2 1B | $0.100 | $0.100 | 128K |
| Mistral Ministral 8B | $0.100 | $0.100 | 128K |
| Amazon Nova Micro | $0.035 | $0.140 | 128K |
| Meta Llama 3.2 3B | $0.150 | $0.150 | 128K |
| Mistral Pixtral 12B | $0.150 | $0.150 | 128K |
| Meta Llama 3.2 11B | $0.160 | $0.160 | 128K |
| Google Gemma 2 9B | $0.200 | $0.200 | 8K |
| Alibaba Qwen 3 14B | $0.080 | $0.240 | 41K |
| Amazon Nova Lite | $0.060 | $0.240 | 300K |
| Mistral Devstral Small | $0.070 | $0.280 | 128K |
| Alibaba Qwen 3 30B | $0.100 | $0.300 | 41K |
| Alibaba Qwen 3 32B | $0.100 | $0.300 | 41K |
| Google Gemini 2.0 Flash Lite | $0.075 | $0.300 | 1.0M |
| Meta Llama 4 Scout | $0.100 | $0.300 | 131K |
| Mistral Mistral Small | $0.100 | $0.300 | 32K |
| OpenAI GPT 4.1 Nano | $0.100 | $0.400 | 1.0M |
| Xai Grok 3 Mini | $0.300 | $0.500 | 131K |
| Alibaba Qwen 3 235B | $0.200 | $0.600 | 41K |
| Cohere Command R | $0.150 | $0.600 | 128K |
| Google Gemini 2.0 Flash | $0.150 | $0.600 | 1.0M |
| Meta Llama 4 Maverick | $0.200 | $0.600 | 131K |
| OpenAI GPT 4o Mini | $0.150 | $0.600 | 128K |
| Meta Llama 3.1 70B | $0.720 | $0.720 | 128K |
| Meta Llama 3.2 90B | $0.720 | $0.720 | 128K |
| Meta Llama 3.3 70B | $0.720 | $0.720 | 128K |
| Meta Llama 3 70B | $0.590 | $0.790 | 8K |
| Mistral Mistral Saba 24B | $0.790 | $0.790 | 33K |
| DeepSeek DeepSeek | $0.900 | $0.900 | 128K |
| Mistral Codestral | $0.300 | $0.900 | 256K |
| DeepSeek DeepSeek R1 Distill Llama 70B | $0.750 | $0.990 | 131K |
| Inception Mercury Coder Small | $0.250 | $1.00 | 32K |
| Perplexity Sonar | $1.00 | $1.00 | 127K |
| Zai Glm 4.5 Air | $0.200 | $1.10 | 128K |
| Mistral Mixtral 8X22B Instruct | $1.20 | $1.20 | 66K |
| Morph Morph V3 Fast | $0.800 | $1.20 | 33K |
| Anthropic Claude 3 Haiku | $0.250 | $1.25 | 200K |
| Mistral Magistral Small | $0.500 | $1.50 | 128K |
| OpenAI GPT 3.5 Turbo | $0.500 | $1.50 | 16K |
| Alibaba Qwen3 Coder | $0.400 | $1.60 | 262K |
| OpenAI GPT 4.1 Mini | $0.400 | $1.60 | 1.0M |
| Zai Glm 4.6 | $0.450 | $1.80 | 200K |
| Morph Morph V3 Large | $0.900 | $1.90 | 33K |
| OpenAI GPT 3.5 Turbo Instruct | $1.50 | $2.00 | 8K |
| DeepSeek DeepSeek R1 | $0.550 | $2.19 | 128K |
| Moonshotai Kimi K2 | $0.550 | $2.20 | 131K |
| Zai Glm 4.5 | $0.600 | $2.20 | 131K |
| Google Gemini 2.5 Flash | $0.300 | $2.50 | 1.0M |
| Amazon Nova Pro | $0.800 | $3.20 | 300K |
| Anthropic Claude 3.5 Haiku | $0.800 | $4.00 | 200K |
| Xai Grok 3 Mini Fast | $0.600 | $4.00 | 131K |
| OpenAI o3 Mini | $1.10 | $4.40 | 200K |
| OpenAI o4 Mini | $1.10 | $4.40 | 200K |
| Anthropic Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| Mistral Magistral Medium | $2.00 | $5.00 | 128K |
| Perplexity Sonar Reasoning | $1.00 | $5.00 | 127K |
| Mistral Mistral Large | $2.00 | $6.00 | 32K |
| Mistral Pixtral Large | $2.00 | $6.00 | 128K |
| OpenAI GPT 4.1 | $2.00 | $8.00 | 1.0M |
| OpenAI o3 | $2.00 | $8.00 | 200K |
| Perplexity Sonar Reasoning Pro | $2.00 | $8.00 | 127K |
| Cohere Command A | $2.50 | $10.00 | 256K |
| Cohere Command R Plus | $2.50 | $10.00 | 128K |
| Google Gemini 2.5 Pro | $2.50 | $10.00 | 1.0M |
| OpenAI GPT 4o | $2.50 | $10.00 | 128K |
| Xai Grok 2 | $2.00 | $10.00 | 131K |
| Xai Grok 2 Vision | $2.00 | $10.00 | 33K |
| Anthropic Claude 3.5 Sonnet | $3.00 | $15.00 | 200K |
| Anthropic Claude 3.7 Sonnet | $3.00 | $15.00 | 200K |
| Anthropic Claude 4 Sonnet | $3.00 | $15.00 | 200K |
| Anthropic Claude Sonnet 4.5 | $3.00 | $15.00 | 1.0M |
| Perplexity Sonar Pro | $3.00 | $15.00 | 200K |
| Vercel V0 1.0 Md | $3.00 | $15.00 | 128K |
| Anthropic Claude Opus 4.1 | $15.00 | $75.00 | 200K |
Frequently asked questions
How much does Vercel AI Gateway cost per 1M tokens?
Vercel AI Gateway pricing starts at $0.040 per 1 million output tokens (Mistral Ministral 3B) across 90 models, with its current flagship Anthropic Claude 3 Opus at $75.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Vercel AI Gateway model?
Mistral Ministral 3B is the cheapest Vercel AI Gateway model on output tokens at $0.040 per 1M, while Amazon Titan Embed Text is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Vercel AI Gateway context window?
Google Gemini 2.0 Flash has the largest context window in the Vercel AI Gateway catalog at 1.0M input tokens, with up to 8K output tokens per response.
How do I calculate my actual Vercel AI Gateway bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator