Gemini on Vertex AI Pricing
Gemini on Vertex AI charges from $0.250 per 1 million output tokens, depending on the model. GPT OSS 20B is the cheapest at $0.250/1M output and $0.070/1M input; its current flagship Nano Banana Pro costs $120.00/1M output. The largest context window is 2.0M tokens (Xai Grok 4.1 Fast Non Reasoning). Prices are USD, current as of August 11, 2026.
Vertex AI is the enterprise surface for Gemini — same models as the Gemini API, billed through Google Cloud with IAM, VPC controls and committed-use discounts. Long-context models bill at a higher rate above their tier threshold.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.070/1M
GPT OSS 20B
Cheapest output
$0.250/1M
GPT OSS 20B
Longest context
2.0M
Xai Grok 4.1 Fast Non Reasoning
Models priced
54
Avg $12.84/1M output
Pricing by model
All prices in USD per 1 million tokens. All 54 priced models, cheapest first. The batch column is the asynchronous batch-API rate, 50% below the standard output rate on the 6 models that publish one.
| Model | Input / 1M | Output / 1M | Batch out / 1M | Context |
|---|---|---|---|---|
| GPT OSS 20B ReasoningTools | $0.070 | $0.250 | — | 131K |
| Gemini 2.0 Flash Lite | $0.075 | $0.300 | — | 1.0M |
| Gemini 2.0 Flash Lite 001 | $0.075 | $0.300 | — | 1.0M |
| GPT OSS 120B ReasoningTools | $0.090 | $0.360 | — | 131K |
| Gemini 2.0 Flash | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.5 Flash Lite Preview 06.17 | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.5 Flash Lite Preview 09 2025 | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.5 Flash-Lite VisionReasoningToolsCache | $0.100 | $0.400 | — | 1.0M |
| Xai Grok 4.1 Fast Non Reasoning | $0.200 | $0.500 | — | 2.0M |
| Xai Grok 4.1 Fast Reasoning | $0.200 | $0.500 | — | 2.0M |
| Gemini 2.0 Flash 001 | $0.150 | $0.600 | — | 1.0M |
| Llama 3.3 70B Instruct Tools | $0.720 | $0.720 | — | 128K |
| Qwen3 235B A22B Instruct ReasoningTools | $0.220 | $0.880 | — | 262K |
| Llama 4 Maverick 17B 128E Instruct VisionTools | $0.350 | $1.15 | — | 524K |
| Gemini 3.1 Flash Lite VisionReasoningToolsCache | $0.250 | $1.50 | $0.750−50% | 1.0M |
| Gemini 3.1 Flash Lite Preview VisionReasoningToolsCache | $0.250 | $1.50 | — | 1.0M |
| Gemini Flash-Lite Latest VisionReasoningToolsCache | $0.250 | $1.50 | — | 1.0M |
| DeepSeek V3.2 ReasoningToolsCache | $0.560 | $1.68 | — | 164K |
| DeepSeek V3.1 ReasoningTools | $0.600 | $1.70 | — | 164K |
| GLM-4.7 ReasoningTools | $0.600 | $2.20 | — | 200K |
| Gemini 2.5 Flash VisionReasoningToolsCache | $0.300 | $2.50 | — | 1.0M |
| Gemini 2.5 Flash Preview 09 2025 | $0.300 | $2.50 | — | 1.0M |
| Gemini 3.5 Flash Lite VisionReasoningToolsCache | $0.300 | $2.50 | $1.25−50% | 1.0M |
| Gemini Robotics Er 1.5 Preview | $0.300 | $2.50 | — | 1.0M |
| Kimi K2 Thinking ReasoningTools | $0.600 | $2.50 | — | 262K |
| Gemini 3 Flash Preview VisionReasoningToolsCache | $0.500 | $3.00 | — | 1.0M |
| GLM-5 ReasoningToolsCache | $1.00 | $3.20 | — | 203K |
| Claude Haiku 4.5 VisionReasoningToolsCache | $1.00 | $5.00 | — | 200K |
| Xai Grok 4.20 Non Reasoning | $2.00 | $6.00 | — | 2.0M |
| Xai Grok 4.20 Reasoning | $2.00 | $6.00 | — | 2.0M |
| Gemini 3.6 Flash VisionReasoningToolsCache | $1.50 | $7.50 | $3.75−50% | 1.0M |
| Gemini 3.5 Flash VisionReasoningToolsCache | $1.50 | $9.00 | — | 1.0M |
| Gemini Flash Latest VisionReasoningToolsCache | $1.50 | $9.00 | — | 1.0M |
| Gemini Omni Flash Preview | $1.50 | $9.00 | — | 1.0M |
| Claude Sonnet 5 VisionReasoningToolsCache | $2.00 | $10.00 | — | 1.0M |
| Gemini 2.5 Computer Use Preview 10 2025 | $1.25 | $10.00 | — | 128K |
| Gemini 2.5 Pro VisionReasoningToolsCache | $1.25 | $10.00 | — | 1.0M |
| Gemini 2.5 Pro Preview Tts | $1.25 | $10.00 | — | 1.0M |
| Gemini 3 Pro Preview | $2.00 | $12.00 | $6.00−50% | 1.0M |
| Gemini 3.1 Pro Preview VisionReasoningToolsCache | $2.00 | $12.00 | $6.00−50% | 1.0M |
| Gemini 3.1 Pro Preview Custom Tools VisionReasoningToolsCache | $2.00 | $12.00 | $6.00−50% | 1.0M |
| Claude Sonnet 4 VisionReasoningToolsCache | $3.00 | $15.00 | — | 200K |
| Claude Sonnet 4.5 VisionReasoningToolsCache | $3.00 | $15.00 | — | 200K |
| Claude Sonnet 4.6 VisionReasoningToolsCache | $3.00 | $15.00 | — | 1.0M |
| Claude Opus 4.5 VisionReasoningToolsCache | $5.00 | $25.00 | — | 200K |
| Claude Opus 4.6 VisionReasoningToolsCache | $5.00 | $25.00 | — | 1.0M |
| Claude Opus 4.7 VisionReasoningToolsCache | $5.00 | $25.00 | — | 1.0M |
| Claude Opus 4.8 VisionReasoningToolsCache | $5.00 | $25.00 | — | 1.0M |
| Claude Opus 5 VisionReasoningToolsCache | $5.00 | $25.00 | — | 1.0M |
| Nano Banana Vision | $0.300 | $30.00 | — | 33K |
| Nano Banana 2 VisionReasoning | $0.500 | $60.00 | — | 131K |
| Claude Opus 4 VisionReasoningToolsCache | $15.00 | $75.00 | — | 200K |
| Claude Opus 4.1 VisionReasoningToolsCache | $15.00 | $75.00 | — | 200K |
| Nano Banana Pro VisionReasoning | $2.00 | $120.00 | — | 66K |
A dash in the batch column means our catalog carries no published batch rate for that model — not that the model has no batch tier. Batch coverage in the upstream pricing data is uneven.
Frequently asked questions
How much does Gemini on Vertex AI cost per 1M tokens?
Gemini on Vertex AI pricing starts at $0.250 per 1 million output tokens (GPT OSS 20B) across 54 models, with its current flagship Nano Banana Pro at $120.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Gemini on Vertex AI model?
GPT OSS 20B is the cheapest Gemini on Vertex AI model on both axes — $0.070 per 1M input tokens and $0.250 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Gemini on Vertex AI context window?
Xai Grok 4.1 Fast Non Reasoning has the largest context window in the Gemini on Vertex AI catalog at 2.0M input tokens, with up to 2.0M output tokens per response.
Does Gemini on Vertex AI support prompt caching?
Yes. Gemini 2.5 Flash-Lite reads cached input at $0.010 per 1M tokens, against $0.100 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Does Gemini on Vertex AI offer batch pricing?
Yes — 6 Gemini on Vertex AI models in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Gemini 3.6 Flash bills $3.75 per 1M output in batch against $7.50 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.
Which Gemini on Vertex AI models support reasoning or vision?
The Gemini on Vertex AI catalog includes 35 reasoning models, 29 vision models, 35 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Gemini on Vertex AI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator