API pricing

Vercel AI Gateway Pricing

Vercel AI Gateway charges from $0.100 per 1 million output tokens, depending on the model. Ministral 3B is the cheapest at $0.100/1M output and $0.100/1M input; its current flagship GLM 5.3 costs $4.40/1M output. The largest context window is 2.0M tokens (Grok 4.20 Beta Non-Reasoning). Prices are USD, current as of August 31, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.030/1M

Qwen 3.7 Flash

Cheapest output

$0.100/1M

Ministral 3B

Longest context

2.0M

Grok 4.20 Beta Non-Reasoning

Models priced

229

Avg $8.87/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 229 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Ministral 3B
VisionTools
$0.100$0.100128K
Qwen 3.7 Flash
VisionReasoningToolsCache
$0.030$0.130991K
Nova Micro
Tools
$0.035$0.140128K
Ministral 8B
VisionTools
$0.150$0.150128K
Mistral Nemo 12B
VisionTools
$0.150$0.150128K
Pixtral 12B 2409
VisionTools
$0.150$0.150128K
DeepSeek V4 Flash 0731
ReasoningToolsCache
$0.076$0.1531.0M
Tencent Hy-MT2-Lite$0.044$0.1778K
Ling 3.0 Flash
ReasoningToolsCache
$0.060$0.180256K
GPT OSS 20B
ReasoningTools
$0.050$0.200131K
GPT OSS Safeguard 20B
ReasoningTools
$0.070$0.200128K
Laguna S 2.1
ReasoningToolsCache
$0.100$0.2001.0M
Ministral 14B
VisionTools
$0.200$0.200256K
Muse Spark 1.2 Contributor
VisionReasoningToolsCache
$0.100$0.2001.0M
Nemotron 3.5 Lightning 30B
ReasoningToolsCache
$0.050$0.200262K
Llama 3.1 8B Instruct
Tools
$0.220$0.220128K
Nvidia Nemotron Nano 9B V2
ReasoningTools
$0.060$0.230131K
Nemotron 3 Nano 30B A3B
ReasoningTools
$0.050$0.240262K
Nova Lite
VisionTools
$0.060$0.240300K
Qwen3-14B
ReasoningTools
$0.120$0.24041K
DeepSeek V4 Flash
ReasoningToolsCache
$0.130$0.2601.0M
MiMo M2.5
VisionReasoningToolsCache
$0.140$0.2801.1M
Tencent Hy-MT2-Plus$0.074$0.2958K
Tencent Hy-MT2-Pro$0.074$0.2958K
Devstral Small 2
VisionTools
$0.100$0.300256K
Mistral Small
VisionTools
$0.100$0.30032K
StepFun 3.5 Flash
VisionReasoningToolsCache
$0.090$0.300262K
Gemini 2.5 Flash Lite
VisionReasoningToolsCache
$0.100$0.4001.0M
Gemma 4 31B IT
VisionReasoningTools
$0.140$0.400262K
GLM 4.7 Flash
ReasoningTools
$0.070$0.400200K
GLM 4.7 FlashX
ReasoningToolsCache
$0.060$0.400200K
GPT-4.1 nano
VisionToolsCache
$0.100$0.4001.0M
GPT-5 nano
VisionReasoningToolsCache
$0.050$0.400400K
Qwen 3.5 Flash
VisionReasoningToolsCache
$0.100$0.4001.0M
Qwen 3.8 Flash Next
VisionReasoningToolsCache
$0.120$0.4001.0M
DeepSeek V3.2
ToolsCache
$0.280$0.420128K
Qwen 3.8 Flash
VisionReasoningToolsCache
$0.160$0.470991K
GLM 5.3 Flash
VisionReasoningToolsCache
$0.150$0.5001.0M
GPT OSS 120B
ReasoningTools
$0.100$0.500131K
Grok 4.1 Fast Non-Reasoning
VisionToolsCache
$0.200$0.5001.0M
Grok 4.1 Fast Reasoning
VisionReasoningToolsCache
$0.200$0.5001.0M
Qwen3-30B-A3B
ReasoningTools
$0.120$0.50041K
Hy3
ReasoningToolsCache
$0.140$0.580262K
Google Gemma 4 26B A4B
VisionReasoningToolsCache
$0.150$0.600262K
GPT OSS Safeguard 120B
ReasoningTools
$0.150$0.600128K
GPT-4o mini
VisionToolsCache
$0.150$0.600128K
Kat Coder Air V2.5
VisionReasoningToolsCache
$0.150$0.600256K
Nvidia Nemotron Nano 12B V2 VL
VisionReasoningTools
$0.200$0.600131K
Qwen 3 Coder 30B A3B Instruct
Tools
$0.150$0.600262K
Qwen 3 32B
ReasoningTools
$0.160$0.640128K
NVIDIA Nemotron 3 Super 120B A12B
ReasoningTools
$0.150$0.650256K
DeepSeek V4 Flash Vision Exp
VisionReasoningToolsCache
$0.220$0.6601.0M
Llama 4 Scout 17B Instruct
VisionTools
$0.170$0.660128K
Llama 3.1 70B Instruct
Tools
$0.720$0.720128K
Llama 3.3 70B Instruct
Tools
$0.720$0.720128K
Mercury 2
ReasoningToolsCache
$0.250$0.750128K
GPT-4.1 nano (Fast)
VisionToolsCache
$0.200$0.8001.0M
MiMo V2.5 Pro
ReasoningToolsCache
$0.435$0.8701.1M
Qwen3 235B A22B
ReasoningTools
$0.220$0.880262K
Mistral Codestral
Tools
$0.300$0.900128K
Trinity Large Thinking
ReasoningTools
$0.250$0.900262K
DeepSeek V3.1
ReasoningToolsCache
$0.250$0.950164K
Llama 4 Maverick 17B Instruct
VisionTools
$0.240$0.970128K
DeepSeek V3.1 Terminus
ReasoningToolsCache
$0.270$1.00131K
GPT-4o mini (Fast)
VisionToolsCache
$0.250$1.00128K
Mercury Coder Small Beta
Tools
$0.250$1.0032K
GLM 4.5 Air
ReasoningToolsCache
$0.200$1.10128K
DeepSeek V3 0324
ToolsCache
$0.270$1.12164K
Step 3.7 Flash
VisionReasoningToolsCache
$0.200$1.15256K
GPT 5.6 Luna
VisionReasoningToolsCache
$0.200$1.201.1M
Inkling Small
VisionReasoningToolsCache
$0.500$1.201.0M
Kat Coder Pro V2
ReasoningToolsCache
$0.300$1.20256K
KAT-Coder-Pro V1
ToolsCache
$0.300$1.20256K
MiniMax M2
ReasoningToolsCache
$0.300$1.20205K
MiniMax M2.1
ReasoningToolsCache
$0.300$1.20205K
MiniMax M2.5
ReasoningToolsCache
$0.300$1.20205K
MiniMax M2.7
ReasoningToolsCache
$0.300$1.20205K
MiniMax M3
VisionReasoningToolsCache
$0.300$1.20512K
Morph V3 Fast$0.800$1.2082K
GPT 5.5 Pro
VisionReasoningTools
$30.00$180.001.0M

Frequently asked questions

How much does Vercel AI Gateway cost per 1M tokens?

Vercel AI Gateway pricing starts at $0.100 per 1 million output tokens (Ministral 3B) across 229 models, with its current flagship GLM 5.3 at $4.40. Older premium models in the catalog list higher. Rates current as of August 31, 2026.

What is the cheapest Vercel AI Gateway model?

Ministral 3B is the cheapest Vercel AI Gateway model on output tokens at $0.100 per 1M, while Qwen 3.7 Flash is cheapest on input at $0.030 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Vercel AI Gateway context window?

Grok 4.20 Beta Non-Reasoning has the largest context window in the Vercel AI Gateway catalog at 2.0M input tokens, with up to 2.0M output tokens per response.

Does Vercel AI Gateway support prompt caching?

Yes. Qwen 3.5 Flash reads cached input at $0.0010 per 1M tokens, against $0.100 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Vercel AI Gateway models support reasoning or vision?

The Vercel AI Gateway catalog includes 172 reasoning models, 150 vision models, 218 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Vercel AI Gateway bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator