Vercel AI Gateway Pricing
Vercel AI Gateway charges between — and $75.00 per 1 million output tokens, depending on the model. Amazon Titan Embed Text is the cheapest at —/1M output and $0.020/1M input; Anthropic Claude 3 Opus is the most expensive at $75.00/1M output. The largest context window is 1.0M tokens (Google Gemini 2.0 Flash). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Amazon Titan Embed Text
Cheapest output
—/1M
Amazon Titan Embed Text
Longest context
1.0M
Google Gemini 2.0 Flash
Models priced
90
Avg $8.63/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 90 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Amazon Titan Embed Text | $0.020 | — | 0 |
| Cohere Embed V4.0 | $0.120 | — | 0 |
| Mistral Codestral Embed | $0.150 | — | 0 |
| Mistral Mistral Embed | $0.100 | — | 0 |
| Mistral Ministral 3B | $0.040 | $0.040 | 128K |
| Meta Llama 3 8B | $0.050 | $0.080 | 8K |
| Meta Llama 3.1 8B | $0.050 | $0.080 | 131K |
| Meta Llama 3.2 1B | $0.100 | $0.100 | 128K |
| Mistral Ministral 8B | $0.100 | $0.100 | 128K |
| Amazon Nova Micro | $0.035 | $0.140 | 128K |
| Meta Llama 3.2 3B | $0.150 | $0.150 | 128K |
| Mistral Pixtral 12B | $0.150 | $0.150 | 128K |
| Meta Llama 3.2 11B | $0.160 | $0.160 | 128K |
| Google Gemma 2 9B | $0.200 | $0.200 | 8K |
| Alibaba Qwen 3 14B | $0.080 | $0.240 | 41K |
| Amazon Nova Lite | $0.060 | $0.240 | 300K |
| Mistral Devstral Small | $0.070 | $0.280 | 128K |
| Alibaba Qwen 3 30B | $0.100 | $0.300 | 41K |
| Alibaba Qwen 3 32B | $0.100 | $0.300 | 41K |
| Google Gemini 2.0 Flash Lite | $0.075 | $0.300 | 1.0M |
| Meta Llama 4 Scout | $0.100 | $0.300 | 131K |
| Mistral Mistral Small | $0.100 | $0.300 | 32K |
| OpenAI GPT 4.1 Nano | $0.100 | $0.400 | 1.0M |
| Xai Grok 3 Mini | $0.300 | $0.500 | 131K |
| Alibaba Qwen 3 235B | $0.200 | $0.600 | 41K |
| Cohere Command R | $0.150 | $0.600 | 128K |
| Google Gemini 2.0 Flash | $0.150 | $0.600 | 1.0M |
| Meta Llama 4 Maverick | $0.200 | $0.600 | 131K |
| OpenAI GPT 4o Mini | $0.150 | $0.600 | 128K |
| Meta Llama 3.1 70B | $0.720 | $0.720 | 128K |
| Meta Llama 3.2 90B | $0.720 | $0.720 | 128K |
| Meta Llama 3.3 70B | $0.720 | $0.720 | 128K |
| Meta Llama 3 70B | $0.590 | $0.790 | 8K |
| Mistral Mistral Saba 24B | $0.790 | $0.790 | 33K |
| DeepSeek DeepSeek | $0.900 | $0.900 | 128K |
| Mistral Codestral | $0.300 | $0.900 | 256K |
| DeepSeek DeepSeek R1 Distill Llama 70B | $0.750 | $0.990 | 131K |
| Inception Mercury Coder Small | $0.250 | $1.00 | 32K |
| Perplexity Sonar | $1.00 | $1.00 | 127K |
| Zai Glm 4.5 Air | $0.200 | $1.10 | 128K |
| Mistral Mixtral 8X22B Instruct | $1.20 | $1.20 | 66K |
| Morph Morph V3 Fast | $0.800 | $1.20 | 33K |
| Anthropic Claude 3 Haiku | $0.250 | $1.25 | 200K |
| Mistral Magistral Small | $0.500 | $1.50 | 128K |
| OpenAI GPT 3.5 Turbo | $0.500 | $1.50 | 16K |
| Alibaba Qwen3 Coder | $0.400 | $1.60 | 262K |
| OpenAI GPT 4.1 Mini | $0.400 | $1.60 | 1.0M |
| Zai Glm 4.6 | $0.450 | $1.80 | 200K |
| Morph Morph V3 Large | $0.900 | $1.90 | 33K |
| OpenAI GPT 3.5 Turbo Instruct | $1.50 | $2.00 | 8K |
| DeepSeek DeepSeek R1 | $0.550 | $2.19 | 128K |
| Moonshotai Kimi K2 | $0.550 | $2.20 | 131K |
| Zai Glm 4.5 | $0.600 | $2.20 | 131K |
| Google Gemini 2.5 Flash | $0.300 | $2.50 | 1.0M |
| Amazon Nova Pro | $0.800 | $3.20 | 300K |
| Anthropic Claude 3.5 Haiku | $0.800 | $4.00 | 200K |
| Xai Grok 3 Mini Fast | $0.600 | $4.00 | 131K |
| OpenAI o3 Mini | $1.10 | $4.40 | 200K |
| OpenAI o4 Mini | $1.10 | $4.40 | 200K |
| Anthropic Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| Mistral Magistral Medium | $2.00 | $5.00 | 128K |
| Perplexity Sonar Reasoning | $1.00 | $5.00 | 127K |
| Mistral Mistral Large | $2.00 | $6.00 | 32K |
| Mistral Pixtral Large | $2.00 | $6.00 | 128K |
| OpenAI GPT 4.1 | $2.00 | $8.00 | 1.0M |
| OpenAI o3 | $2.00 | $8.00 | 200K |
| Perplexity Sonar Reasoning Pro | $2.00 | $8.00 | 127K |
| Cohere Command A | $2.50 | $10.00 | 256K |
| Cohere Command R Plus | $2.50 | $10.00 | 128K |
| Google Gemini 2.5 Pro | $2.50 | $10.00 | 1.0M |
| OpenAI GPT 4o | $2.50 | $10.00 | 128K |
| Xai Grok 2 | $2.00 | $10.00 | 131K |
| Xai Grok 2 Vision | $2.00 | $10.00 | 33K |
| Anthropic Claude 3.5 Sonnet | $3.00 | $15.00 | 200K |
| Anthropic Claude 3.7 Sonnet | $3.00 | $15.00 | 200K |
| Anthropic Claude 4 Sonnet | $3.00 | $15.00 | 200K |
| Anthropic Claude Sonnet 4.5 | $3.00 | $15.00 | 1.0M |
| Perplexity Sonar Pro | $3.00 | $15.00 | 200K |
| Vercel V0 1.0 Md | $3.00 | $15.00 | 128K |
| Anthropic Claude Opus 4.1 | $15.00 | $75.00 | 200K |
Frequently asked questions
How much does Vercel AI Gateway cost per 1M tokens?
Vercel AI Gateway pricing ranges from — to $75.00 per 1 million output tokens across 90 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Vercel AI Gateway model?
Amazon Titan Embed Text is the cheapest Vercel AI Gateway model on both axes — $0.020 per 1M input tokens and — per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Vercel AI Gateway context window?
Google Gemini 2.0 Flash has the largest context window in the Vercel AI Gateway catalog at 1.0M input tokens, with up to 8K output tokens per response.
How do I calculate my actual Vercel AI Gateway bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator