API pricing

Gemini on Vertex AI Pricing

Gemini on Vertex AI charges between $0.250 and $120.00 per 1 million output tokens, depending on the model. GPT OSS 20B is the cheapest at $0.250/1M output and $0.070/1M input; Nano Banana Pro is the most expensive at $120.00/1M output. The largest context window is 2.0M tokens (Xai Grok 4.1 Fast Non Reasoning). Prices are USD, current as of July 20, 2026.

Vertex AI is the enterprise surface for Gemini — same models as the Gemini API, billed through Google Cloud with IAM, VPC controls and committed-use discounts. Long-context models bill at a higher rate above their tier threshold.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.070/1M

GPT OSS 20B

Cheapest output

$0.250/1M

GPT OSS 20B

Longest context

2.0M

Xai Grok 4.1 Fast Non Reasoning

Models priced

52

Avg $12.74/1M output

Pricing by model

All prices in USD per 1 million tokens. All 52 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
GPT OSS 20B
ReasoningTools
$0.070$0.250131K
Gemini 2.0 Flash Lite$0.075$0.3001.0M
Gemini 2.0 Flash Lite 001$0.075$0.3001.0M
GPT OSS 120B
ReasoningTools
$0.090$0.360131K
Gemini 2.0 Flash$0.100$0.4001.0M
Gemini 2.5 Flash Lite Preview 06.17$0.100$0.4001.0M
Gemini 2.5 Flash Lite Preview 09 2025$0.100$0.4001.0M
Gemini 2.5 Flash-Lite
VisionReasoningToolsCache
$0.100$0.4001.0M
Xai Grok 4.1 Fast Non Reasoning$0.200$0.5002.0M
Xai Grok 4.1 Fast Reasoning$0.200$0.5002.0M
Gemini 2.0 Flash 001$0.150$0.6001.0M
Llama 3.3 70B Instruct
Tools
$0.720$0.720128K
Qwen3 235B A22B Instruct
ReasoningTools
$0.220$0.880262K
Llama 4 Maverick 17B 128E Instruct
VisionTools
$0.350$1.15524K
Gemini 3.1 Flash Lite
VisionReasoningToolsCache
$0.250$1.501.0M
Gemini 3.1 Flash Lite Preview
VisionReasoningToolsCache
$0.250$1.501.0M
Gemini Flash-Lite Latest
VisionReasoningToolsCache
$0.250$1.501.0M
DeepSeek V3.2
ReasoningToolsCache
$0.560$1.68164K
DeepSeek V3.1
ReasoningTools
$0.600$1.70164K
GLM-4.7
ReasoningTools
$0.600$2.20200K
Gemini 2.5 Flash
VisionReasoningToolsCache
$0.300$2.501.0M
Gemini 2.5 Flash Preview 09 2025$0.300$2.501.0M
Gemini Robotics Er 1.5 Preview$0.300$2.501.0M
Kimi K2 Thinking
ReasoningTools
$0.600$2.50262K
Gemini 3 Flash Preview
VisionReasoningToolsCache
$0.500$3.001.0M
GLM-5
ReasoningToolsCache
$1.00$3.20203K
Claude Haiku 3.5
VisionToolsCache
$0.800$4.00200K
Claude Haiku 4.5
VisionReasoningToolsCache
$1.00$5.00200K
Xai Grok 4.20 Non Reasoning$2.00$6.002.0M
Xai Grok 4.20 Reasoning$2.00$6.002.0M
Gemini 3.5 Flash
VisionReasoningToolsCache
$1.50$9.001.0M
Gemini Flash Latest
VisionReasoningToolsCache
$1.50$9.001.0M
Gemini Omni Flash Preview$1.50$9.001.0M
Claude Sonnet 5
VisionReasoningToolsCache
$2.00$10.001.0M
Gemini 2.5 Computer Use Preview 10 2025$1.25$10.00128K
Gemini 2.5 Pro
VisionReasoningToolsCache
$1.25$10.001.0M
Gemini 2.5 Pro Preview Tts$1.25$10.001.0M
Gemini 3 Pro Preview$2.00$12.001.0M
Gemini 3.1 Pro Preview
VisionReasoningToolsCache
$2.00$12.001.0M
Gemini 3.1 Pro Preview Custom Tools
VisionReasoningToolsCache
$2.00$12.001.0M
Claude Sonnet 4
VisionReasoningToolsCache
$3.00$15.00200K
Claude Sonnet 4.5
VisionReasoningToolsCache
$3.00$15.00200K
Claude Sonnet 4.6
VisionReasoningToolsCache
$3.00$15.001.0M
Claude Opus 4.5
VisionReasoningToolsCache
$5.00$25.00200K
Claude Opus 4.6
VisionReasoningToolsCache
$5.00$25.001.0M
Claude Opus 4.7
VisionReasoningToolsCache
$5.00$25.001.0M
Claude Opus 4.8
VisionReasoningToolsCache
$5.00$25.001.0M
Nano Banana
Vision
$0.300$30.0033K
Nano Banana 2
VisionReasoning
$0.500$60.00131K
Claude Opus 4
VisionReasoningToolsCache
$15.00$75.00200K
Claude Opus 4.1
VisionReasoningToolsCache
$15.00$75.00200K
Nano Banana Pro
VisionReasoning
$2.00$120.0066K

Frequently asked questions

How much does Gemini on Vertex AI cost per 1M tokens?

Gemini on Vertex AI pricing ranges from $0.250 to $120.00 per 1 million output tokens across 52 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Gemini on Vertex AI model?

GPT OSS 20B is the cheapest Gemini on Vertex AI model on both axes — $0.070 per 1M input tokens and $0.250 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Gemini on Vertex AI context window?

Xai Grok 4.1 Fast Non Reasoning has the largest context window in the Gemini on Vertex AI catalog at 2.0M input tokens, with up to 2.0M output tokens per response.

Does Gemini on Vertex AI support prompt caching?

Yes. Gemini 2.5 Flash-Lite reads cached input at $0.010 per 1M tokens, against $0.100 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Gemini on Vertex AI models support reasoning or vision?

The Gemini on Vertex AI catalog includes 32 reasoning models, 27 vision models, 33 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Gemini on Vertex AI bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator