Google Vertex AI Pricing
Google Vertex AI charges from $0.040 per 1 million output tokens, depending on the model. Gemma 3n 4B is the cheapest at $0.040/1M output and $0.020/1M input; its current flagship Nano Banana 2 Lite costs $30.00/1M output. The largest context window is 2.0M tokens (Gemini 3/3.1 (> 200k context)). Prices are USD, current as of August 11, 2026.
Google prices Gemini per million tokens through the Gemini API, with tiered rates on the long-context models — prompts above the tier threshold bill at a higher rate. Vertex AI carries the same model lineup under GCP billing.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.017/1M
Gemma 3 4B
Cheapest output
$0.040/1M
Gemma 3n 4B
Longest context
2.0M
Gemini 3/3.1 (> 200k context)
Models priced
52
Avg $9.45/1M output
Pricing by model
All prices in USD per 1 million tokens. All 52 priced models, cheapest first. The batch column is the asynchronous batch-API rate, 50% below the standard output rate on the 6 models that publish one.
| Model | Input / 1M | Output / 1M | Batch out / 1M | Context |
|---|---|---|---|---|
| Gemma 3n 4B | $0.020 | $0.040 | — | 33K |
| Gemma 3 4B | $0.017 | $0.068 | — | 96K |
| Gemma 2 9B | $0.030 | $0.090 | — | 8K |
| Gemma 3 12B | $0.030 | $0.100 | — | 131K |
| Gemma 3 27B | $0.090 | $0.160 | — | 131K |
| Gemini 2.0 Flash Lite | $0.075 | $0.300 | — | 1.0M |
| Gemini 2.0 Flash Lite 001 | $0.075 | $0.300 | — | 1.0M |
| Gemini 2.0 Flash | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.0 Flash 001 | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.5 Flash Lite | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.5 Flash Lite Preview 06-17 | $0.100 | $0.400 | — | 1.0M |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.100 | $0.400 | — | 1.0M |
| Gemma 2 27B | $0.650 | $0.650 | — | 8K |
| Gemini Gemma 2 27B It | $0.350 | $1.05 | — | 8K |
| Gemini Gemma 2 9B It | $0.350 | $1.05 | — | 8K |
| Gemini 3.1 Flash Lite VisionReasoningToolsCache | $0.250 | $1.50 | $0.750−50% | 1.0M |
| Gemini 3.1 Flash Lite Preview VisionReasoningToolsCache | $0.250 | $1.50 | — | 1.0M |
| Gemini Flash-Lite Latest VisionReasoningToolsCache | $0.250 | $1.50 | — | 1.0M |
| Gemini 2.5 Flash | $0.300 | $2.50 | — | 1.0M |
| Gemini 2.5 Flash Image (Nano Banana) | $0.300 | $2.50 | — | 33K |
| Gemini 2.5 Flash Image Preview (Nano Banana) | $0.300 | $2.50 | — | 33K |
| Gemini 2.5 Flash Native Audio Latest | $0.300 | $2.50 | — | 1.0M |
| Gemini 2.5 Flash Native Audio Preview 09 2025 | $0.300 | $2.50 | — | 1.0M |
| Gemini 2.5 Flash Native Audio Preview 12 2025 | $0.300 | $2.50 | — | 1.0M |
| Gemini 2.5 Flash Preview 09-2025 | $0.300 | $2.50 | — | 1.0M |
| Gemini 3.5 Flash Lite VisionReasoningToolsCache | $0.300 | $2.50 | $1.25−50% | 1.0M |
| Gemini Robotics Er 1.5 Preview | $0.300 | $2.50 | — | 1.0M |
| Gemini 3 Flash Preview | $0.500 | $3.00 | — | 1.0M |
| Gemini 3.1 Flash Live Preview VisionReasoningTools | $0.750 | $4.50 | — | 131K |
| Gemini Robotics-ER 1.6 Preview VisionReasoningTools | $1.00 | $5.00 | — | 131K |
| Gemini 3.6 Flash VisionReasoningToolsCache | $1.50 | $7.50 | $3.75−50% | 1.0M |
| Gemini 3.5 Flash VisionReasoningToolsCache | $1.50 | $9.00 | — | 1.0M |
| Gemini Flash Latest VisionReasoningToolsCache | $1.50 | $9.00 | — | 1.0M |
| Gemini Omni Flash Preview | $1.50 | $9.00 | — | 1.0M |
| Gemini 2.5 Computer Use Preview 10-2025 VisionReasoningTools | $1.25 | $10.00 | — | 131K |
| Gemini 2.5 Pro | $1.25 | $10.00 | — | 1.0M |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 | — | 1.0M |
| Gemini 2.5 Pro Preview Tts | $1.25 | $10.00 | — | 1.0M |
| Gemini Pro Latest | $1.25 | $10.00 | — | 1.0M |
| Gemini Robotics Er 2 Preview | $2.00 | $10.00 | — | 131K |
| Deep Research Max Preview (Apr-21-2026) VisionReasoningToolsCache | $2.00 | $12.00 | — | 131K |
| Deep Research Preview (Apr-21-2026) VisionReasoningToolsCache | $2.00 | $12.00 | — | 131K |
| Gemini 3 Pro Preview | $2.00 | $12.00 | $6.00−50% | 1.0M |
| Gemini 3.1 Pro Preview VisionReasoningToolsCache | $2.00 | $12.00 | $6.00−50% | 1.0M |
| Gemini 3.1 Pro Preview Custom Tools VisionReasoningToolsCache | $2.00 | $12.00 | $6.00−50% | 1.0M |
| Gemini 3/3.1 (≤ 200k context) | $2.00 | $12.00 | — | 200K |
| Gemini 3/3.1 (> 200k context) | $4.00 | $18.00 | — | 2.0M |
| Gemini 3.5 Live Translate Preview | $3.50 | $21.00 | — | 16K |
| Nano Banana VisionReasoningCache | $0.300 | $30.00 | — | 33K |
| Nano Banana 2 Lite VisionReasoningTools | $0.250 | $30.00 | — | 66K |
| Nano Banana 2 VisionReasoning | $0.500 | $60.00 | — | 66K |
| Nano Banana Pro VisionReasoning | $2.00 | $120.00 | — | 131K |
A dash in the batch column means our catalog carries no published batch rate for that model — not that the model has no batch tier. Batch coverage in the upstream pricing data is uneven.
Frequently asked questions
How much does Google Vertex AI cost per 1M tokens?
Google Vertex AI pricing starts at $0.040 per 1 million output tokens (Gemma 3n 4B) across 52 models, with its current flagship Nano Banana 2 Lite at $30.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Google Vertex AI model?
Gemma 3n 4B is the cheapest Google Vertex AI model on output tokens at $0.040 per 1M, while Gemma 3 4B is cheapest on input at $0.017 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Google Vertex AI context window?
Gemini 3/3.1 (> 200k context) has the largest context window in the Google Vertex AI catalog at 2.0M input tokens, with up to 16K output tokens per response.
Does Google Vertex AI support prompt caching?
Yes. Gemini 3.1 Flash Lite reads cached input at $0.025 per 1M tokens, against $0.250 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Does Google Vertex AI offer batch pricing?
Yes — 6 Google Vertex AI models in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Gemini 3.6 Flash bills $3.75 per 1M output in batch against $7.50 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.
Which Google Vertex AI models support reasoning or vision?
The Google Vertex AI catalog includes 18 reasoning models, 18 vision models, 15 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Google Vertex AI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Head-to-head comparisons
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator