API pricing

Google Vertex AI Pricing

Google Vertex AI charges between $0.040 and $60.00 per 1 million output tokens, depending on the model. Gemma 3n 4B is the cheapest at $0.040/1M output and $0.020/1M input; Nano Banana 2 is the most expensive at $60.00/1M output. The largest context window is 2.0M tokens (Gemini 3/3.1 (> 200k context)). Prices are USD, current as of July 20, 2026.

Google prices Gemini per million tokens through the Gemini API, with tiered rates on the long-context models — prompts above the tier threshold bill at a higher rate. Vertex AI carries the same model lineup under GCP billing.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.017/1M

Gemma 3 4B

Cheapest output

$0.040/1M

Gemma 3n 4B

Longest context

2.0M

Gemini 3/3.1 (> 200k context)

Models priced

44

Avg $8.91/1M output

Pricing by model

All prices in USD per 1 million tokens. All 44 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Gemma 3n 4B$0.020$0.04033K
Gemma 3 4B$0.017$0.06896K
Gemma 2 9B$0.030$0.0908K
Gemma 3 12B$0.030$0.100131K
Gemma 3 27B$0.090$0.160131K
Gemini 2.0 Flash Lite$0.075$0.3001.0M
Gemini 2.0 Flash Lite 001$0.075$0.3001.0M
Gemini 2.0 Flash$0.100$0.4001.0M
Gemini 2.0 Flash 001$0.100$0.4001.0M
Gemini 2.5 Flash Lite$0.100$0.4001.0M
Gemini 2.5 Flash Lite Preview 06-17$0.100$0.4001.0M
Gemini 2.5 Flash Lite Preview 09-2025$0.100$0.4001.0M
Gemma 2 27B$0.650$0.6508K
Gemini Gemma 2 27B It$0.350$1.058K
Gemini Gemma 2 9B It$0.350$1.058K
Gemini 3.1 Flash Lite
VisionReasoningToolsCache
$0.250$1.501.0M
Gemini 3.1 Flash Lite Preview
VisionReasoningToolsCache
$0.250$1.501.0M
Gemini Flash-Lite Latest
VisionReasoningToolsCache
$0.250$1.501.0M
Gemini 2.5 Flash$0.300$2.501.0M
Gemini 2.5 Flash Image (Nano Banana)$0.300$2.5033K
Gemini 2.5 Flash Image Preview (Nano Banana)$0.300$2.5033K
Gemini 2.5 Flash Native Audio Latest$0.300$2.501.0M
Gemini 2.5 Flash Native Audio Preview 09 2025$0.300$2.501.0M
Gemini 2.5 Flash Native Audio Preview 12 2025$0.300$2.501.0M
Gemini 2.5 Flash Preview 09-2025$0.300$2.501.0M
Gemini Robotics Er 1.5 Preview$0.300$2.501.0M
Gemini 3 Flash Preview$0.500$3.001.0M
Gemini 3.1 Flash Live Preview$0.750$4.50131K
Gemini 3.5 Flash
VisionReasoningToolsCache
$1.50$9.001.0M
Gemini Flash Latest
VisionReasoningToolsCache
$1.50$9.001.0M
Gemini Omni Flash Preview$1.50$9.001.0M
Gemini 2.5 Computer Use Preview 10 2025$1.25$10.00128K
Gemini 2.5 Pro$1.25$10.001.0M
Gemini 2.5 Pro Preview 05-06$1.25$10.001.0M
Gemini 2.5 Pro Preview Tts$1.25$10.001.0M
Gemini Pro Latest$1.25$10.001.0M
Gemini 3 Pro Preview
VisionReasoningToolsCache
$2.00$12.001.0M
Gemini 3.1 Pro Preview
VisionReasoningToolsCache
$2.00$12.001.0M
Gemini 3.1 Pro Preview Custom Tools
VisionReasoningToolsCache
$2.00$12.001.0M
Gemini 3/3.1 (≤ 200k context)$2.00$12.00200K
Gemini 3/3.1 (> 200k context)$4.00$18.002.0M
Nano Banana
VisionReasoningCache
$0.300$30.0033K
Nano Banana 2
VisionReasoning
$0.500$60.0066K
Nano Banana Pro
VisionReasoning
$2.00$120.00131K

Frequently asked questions

How much does Google Vertex AI cost per 1M tokens?

Google Vertex AI pricing ranges from $0.040 to $60.00 per 1 million output tokens across 44 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Google Vertex AI model?

Gemma 3n 4B is the cheapest Google Vertex AI model on output tokens at $0.040 per 1M, while Gemma 3 4B is cheapest on input at $0.017 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Google Vertex AI context window?

Gemini 3/3.1 (> 200k context) has the largest context window in the Google Vertex AI catalog at 2.0M input tokens, with up to 16K output tokens per response.

Does Google Vertex AI support prompt caching?

Yes. Gemini 3.1 Flash Lite reads cached input at $0.025 per 1M tokens, against $0.250 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Google Vertex AI models support reasoning or vision?

The Google Vertex AI catalog includes 11 reasoning models, 11 vision models, 8 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Google Vertex AI bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Head-to-head comparisons

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator