Vercel AI Gateway Pricing
Vercel AI Gateway charges from $0.100 per 1 million output tokens, depending on the model. Ministral 3B is the cheapest at $0.100/1M output and $0.100/1M input; its current flagship GLM 5.3 costs $4.40/1M output. The largest context window is 2.0M tokens (Grok 4.20 Beta Non-Reasoning). Prices are USD, current as of August 31, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.030/1M
Qwen 3.7 Flash
Cheapest output
$0.100/1M
Ministral 3B
Longest context
2.0M
Grok 4.20 Beta Non-Reasoning
Models priced
229
Avg $8.87/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 229 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Ministral 3B VisionTools | $0.100 | $0.100 | 128K |
| Qwen 3.7 Flash VisionReasoningToolsCache | $0.030 | $0.130 | 991K |
| Nova Micro Tools | $0.035 | $0.140 | 128K |
| Ministral 8B VisionTools | $0.150 | $0.150 | 128K |
| Mistral Nemo 12B VisionTools | $0.150 | $0.150 | 128K |
| Pixtral 12B 2409 VisionTools | $0.150 | $0.150 | 128K |
| DeepSeek V4 Flash 0731 ReasoningToolsCache | $0.076 | $0.153 | 1.0M |
| Tencent Hy-MT2-Lite | $0.044 | $0.177 | 8K |
| Ling 3.0 Flash ReasoningToolsCache | $0.060 | $0.180 | 256K |
| GPT OSS 20B ReasoningTools | $0.050 | $0.200 | 131K |
| GPT OSS Safeguard 20B ReasoningTools | $0.070 | $0.200 | 128K |
| Laguna S 2.1 ReasoningToolsCache | $0.100 | $0.200 | 1.0M |
| Ministral 14B VisionTools | $0.200 | $0.200 | 256K |
| Muse Spark 1.2 Contributor VisionReasoningToolsCache | $0.100 | $0.200 | 1.0M |
| Nemotron 3.5 Lightning 30B ReasoningToolsCache | $0.050 | $0.200 | 262K |
| Llama 3.1 8B Instruct Tools | $0.220 | $0.220 | 128K |
| Nvidia Nemotron Nano 9B V2 ReasoningTools | $0.060 | $0.230 | 131K |
| Nemotron 3 Nano 30B A3B ReasoningTools | $0.050 | $0.240 | 262K |
| Nova Lite VisionTools | $0.060 | $0.240 | 300K |
| Qwen3-14B ReasoningTools | $0.120 | $0.240 | 41K |
| DeepSeek V4 Flash ReasoningToolsCache | $0.130 | $0.260 | 1.0M |
| MiMo M2.5 VisionReasoningToolsCache | $0.140 | $0.280 | 1.1M |
| Tencent Hy-MT2-Plus | $0.074 | $0.295 | 8K |
| Tencent Hy-MT2-Pro | $0.074 | $0.295 | 8K |
| Devstral Small 2 VisionTools | $0.100 | $0.300 | 256K |
| Mistral Small VisionTools | $0.100 | $0.300 | 32K |
| StepFun 3.5 Flash VisionReasoningToolsCache | $0.090 | $0.300 | 262K |
| Gemini 2.5 Flash Lite VisionReasoningToolsCache | $0.100 | $0.400 | 1.0M |
| Gemma 4 31B IT VisionReasoningTools | $0.140 | $0.400 | 262K |
| GLM 4.7 Flash ReasoningTools | $0.070 | $0.400 | 200K |
| GLM 4.7 FlashX ReasoningToolsCache | $0.060 | $0.400 | 200K |
| GPT-4.1 nano VisionToolsCache | $0.100 | $0.400 | 1.0M |
| GPT-5 nano VisionReasoningToolsCache | $0.050 | $0.400 | 400K |
| Qwen 3.5 Flash VisionReasoningToolsCache | $0.100 | $0.400 | 1.0M |
| Qwen 3.8 Flash Next VisionReasoningToolsCache | $0.120 | $0.400 | 1.0M |
| DeepSeek V3.2 ToolsCache | $0.280 | $0.420 | 128K |
| Qwen 3.8 Flash VisionReasoningToolsCache | $0.160 | $0.470 | 991K |
| GLM 5.3 Flash VisionReasoningToolsCache | $0.150 | $0.500 | 1.0M |
| GPT OSS 120B ReasoningTools | $0.100 | $0.500 | 131K |
| Grok 4.1 Fast Non-Reasoning VisionToolsCache | $0.200 | $0.500 | 1.0M |
| Grok 4.1 Fast Reasoning VisionReasoningToolsCache | $0.200 | $0.500 | 1.0M |
| Qwen3-30B-A3B ReasoningTools | $0.120 | $0.500 | 41K |
| Hy3 ReasoningToolsCache | $0.140 | $0.580 | 262K |
| Google Gemma 4 26B A4B VisionReasoningToolsCache | $0.150 | $0.600 | 262K |
| GPT OSS Safeguard 120B ReasoningTools | $0.150 | $0.600 | 128K |
| GPT-4o mini VisionToolsCache | $0.150 | $0.600 | 128K |
| Kat Coder Air V2.5 VisionReasoningToolsCache | $0.150 | $0.600 | 256K |
| Nvidia Nemotron Nano 12B V2 VL VisionReasoningTools | $0.200 | $0.600 | 131K |
| Qwen 3 Coder 30B A3B Instruct Tools | $0.150 | $0.600 | 262K |
| Qwen 3 32B ReasoningTools | $0.160 | $0.640 | 128K |
| NVIDIA Nemotron 3 Super 120B A12B ReasoningTools | $0.150 | $0.650 | 256K |
| DeepSeek V4 Flash Vision Exp VisionReasoningToolsCache | $0.220 | $0.660 | 1.0M |
| Llama 4 Scout 17B Instruct VisionTools | $0.170 | $0.660 | 128K |
| Llama 3.1 70B Instruct Tools | $0.720 | $0.720 | 128K |
| Llama 3.3 70B Instruct Tools | $0.720 | $0.720 | 128K |
| Mercury 2 ReasoningToolsCache | $0.250 | $0.750 | 128K |
| GPT-4.1 nano (Fast) VisionToolsCache | $0.200 | $0.800 | 1.0M |
| MiMo V2.5 Pro ReasoningToolsCache | $0.435 | $0.870 | 1.1M |
| Qwen3 235B A22B ReasoningTools | $0.220 | $0.880 | 262K |
| Mistral Codestral Tools | $0.300 | $0.900 | 128K |
| Trinity Large Thinking ReasoningTools | $0.250 | $0.900 | 262K |
| DeepSeek V3.1 ReasoningToolsCache | $0.250 | $0.950 | 164K |
| Llama 4 Maverick 17B Instruct VisionTools | $0.240 | $0.970 | 128K |
| DeepSeek V3.1 Terminus ReasoningToolsCache | $0.270 | $1.00 | 131K |
| GPT-4o mini (Fast) VisionToolsCache | $0.250 | $1.00 | 128K |
| Mercury Coder Small Beta Tools | $0.250 | $1.00 | 32K |
| GLM 4.5 Air ReasoningToolsCache | $0.200 | $1.10 | 128K |
| DeepSeek V3 0324 ToolsCache | $0.270 | $1.12 | 164K |
| Step 3.7 Flash VisionReasoningToolsCache | $0.200 | $1.15 | 256K |
| GPT 5.6 Luna VisionReasoningToolsCache | $0.200 | $1.20 | 1.1M |
| Inkling Small VisionReasoningToolsCache | $0.500 | $1.20 | 1.0M |
| Kat Coder Pro V2 ReasoningToolsCache | $0.300 | $1.20 | 256K |
| KAT-Coder-Pro V1 ToolsCache | $0.300 | $1.20 | 256K |
| MiniMax M2 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax M2.1 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax M2.5 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax M2.7 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax M3 VisionReasoningToolsCache | $0.300 | $1.20 | 512K |
| Morph V3 Fast | $0.800 | $1.20 | 82K |
| GPT 5.5 Pro VisionReasoningTools | $30.00 | $180.00 | 1.0M |
Frequently asked questions
How much does Vercel AI Gateway cost per 1M tokens?
Vercel AI Gateway pricing starts at $0.100 per 1 million output tokens (Ministral 3B) across 229 models, with its current flagship GLM 5.3 at $4.40. Older premium models in the catalog list higher. Rates current as of August 31, 2026.
What is the cheapest Vercel AI Gateway model?
Ministral 3B is the cheapest Vercel AI Gateway model on output tokens at $0.100 per 1M, while Qwen 3.7 Flash is cheapest on input at $0.030 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Vercel AI Gateway context window?
Grok 4.20 Beta Non-Reasoning has the largest context window in the Vercel AI Gateway catalog at 2.0M input tokens, with up to 2.0M output tokens per response.
Does Vercel AI Gateway support prompt caching?
Yes. Qwen 3.5 Flash reads cached input at $0.0010 per 1M tokens, against $0.100 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Vercel AI Gateway models support reasoning or vision?
The Vercel AI Gateway catalog includes 172 reasoning models, 150 vision models, 218 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Vercel AI Gateway bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator