Google Vertex AI Pricing
Google Vertex AI charges between $0.040 and $60.00 per 1 million output tokens, depending on the model. Gemma 3n 4B is the cheapest at $0.040/1M output and $0.020/1M input; Nano Banana 2 is the most expensive at $60.00/1M output. The largest context window is 2.0M tokens (Gemini 3/3.1 (> 200k context)). Prices are USD, current as of July 20, 2026.
Google prices Gemini per million tokens through the Gemini API, with tiered rates on the long-context models — prompts above the tier threshold bill at a higher rate. Vertex AI carries the same model lineup under GCP billing.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.017/1M
Gemma 3 4B
Cheapest output
$0.040/1M
Gemma 3n 4B
Longest context
2.0M
Gemini 3/3.1 (> 200k context)
Models priced
44
Avg $8.91/1M output
Pricing by model
All prices in USD per 1 million tokens. All 44 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Gemma 3n 4B | $0.020 | $0.040 | 33K |
| Gemma 3 4B | $0.017 | $0.068 | 96K |
| Gemma 2 9B | $0.030 | $0.090 | 8K |
| Gemma 3 12B | $0.030 | $0.100 | 131K |
| Gemma 3 27B | $0.090 | $0.160 | 131K |
| Gemini 2.0 Flash Lite | $0.075 | $0.300 | 1.0M |
| Gemini 2.0 Flash Lite 001 | $0.075 | $0.300 | 1.0M |
| Gemini 2.0 Flash | $0.100 | $0.400 | 1.0M |
| Gemini 2.0 Flash 001 | $0.100 | $0.400 | 1.0M |
| Gemini 2.5 Flash Lite | $0.100 | $0.400 | 1.0M |
| Gemini 2.5 Flash Lite Preview 06-17 | $0.100 | $0.400 | 1.0M |
| Gemini 2.5 Flash Lite Preview 09-2025 | $0.100 | $0.400 | 1.0M |
| Gemma 2 27B | $0.650 | $0.650 | 8K |
| Gemini Gemma 2 27B It | $0.350 | $1.05 | 8K |
| Gemini Gemma 2 9B It | $0.350 | $1.05 | 8K |
| Gemini 3.1 Flash Lite VisionReasoningToolsCache | $0.250 | $1.50 | 1.0M |
| Gemini 3.1 Flash Lite Preview VisionReasoningToolsCache | $0.250 | $1.50 | 1.0M |
| Gemini Flash-Lite Latest VisionReasoningToolsCache | $0.250 | $1.50 | 1.0M |
| Gemini 2.5 Flash | $0.300 | $2.50 | 1.0M |
| Gemini 2.5 Flash Image (Nano Banana) | $0.300 | $2.50 | 33K |
| Gemini 2.5 Flash Image Preview (Nano Banana) | $0.300 | $2.50 | 33K |
| Gemini 2.5 Flash Native Audio Latest | $0.300 | $2.50 | 1.0M |
| Gemini 2.5 Flash Native Audio Preview 09 2025 | $0.300 | $2.50 | 1.0M |
| Gemini 2.5 Flash Native Audio Preview 12 2025 | $0.300 | $2.50 | 1.0M |
| Gemini 2.5 Flash Preview 09-2025 | $0.300 | $2.50 | 1.0M |
| Gemini Robotics Er 1.5 Preview | $0.300 | $2.50 | 1.0M |
| Gemini 3 Flash Preview | $0.500 | $3.00 | 1.0M |
| Gemini 3.1 Flash Live Preview | $0.750 | $4.50 | 131K |
| Gemini 3.5 Flash VisionReasoningToolsCache | $1.50 | $9.00 | 1.0M |
| Gemini Flash Latest VisionReasoningToolsCache | $1.50 | $9.00 | 1.0M |
| Gemini Omni Flash Preview | $1.50 | $9.00 | 1.0M |
| Gemini 2.5 Computer Use Preview 10 2025 | $1.25 | $10.00 | 128K |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1.0M |
| Gemini 2.5 Pro Preview 05-06 | $1.25 | $10.00 | 1.0M |
| Gemini 2.5 Pro Preview Tts | $1.25 | $10.00 | 1.0M |
| Gemini Pro Latest | $1.25 | $10.00 | 1.0M |
| Gemini 3 Pro Preview VisionReasoningToolsCache | $2.00 | $12.00 | 1.0M |
| Gemini 3.1 Pro Preview VisionReasoningToolsCache | $2.00 | $12.00 | 1.0M |
| Gemini 3.1 Pro Preview Custom Tools VisionReasoningToolsCache | $2.00 | $12.00 | 1.0M |
| Gemini 3/3.1 (≤ 200k context) | $2.00 | $12.00 | 200K |
| Gemini 3/3.1 (> 200k context) | $4.00 | $18.00 | 2.0M |
| Nano Banana VisionReasoningCache | $0.300 | $30.00 | 33K |
| Nano Banana 2 VisionReasoning | $0.500 | $60.00 | 66K |
| Nano Banana Pro VisionReasoning | $2.00 | $120.00 | 131K |
Frequently asked questions
How much does Google Vertex AI cost per 1M tokens?
Google Vertex AI pricing ranges from $0.040 to $60.00 per 1 million output tokens across 44 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Google Vertex AI model?
Gemma 3n 4B is the cheapest Google Vertex AI model on output tokens at $0.040 per 1M, while Gemma 3 4B is cheapest on input at $0.017 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Google Vertex AI context window?
Gemini 3/3.1 (> 200k context) has the largest context window in the Google Vertex AI catalog at 2.0M input tokens, with up to 16K output tokens per response.
Does Google Vertex AI support prompt caching?
Yes. Gemini 3.1 Flash Lite reads cached input at $0.025 per 1M tokens, against $0.250 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Google Vertex AI models support reasoning or vision?
The Google Vertex AI catalog includes 11 reasoning models, 11 vision models, 8 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Google Vertex AI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Head-to-head comparisons
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator