Google Vertex AI Pricing
Google Vertex AI resells 104 large language models from 6 vendors, billed per million tokens through Google Cloud. Mistral Nemo Latest (Mistral) is the cheapest at $0.150 per 1M output tokens. Anthropic Claude models on Vertex start at $1.25 per 1M output, with Claude Opus 5 at $25.00. Prices are USD, current as of August 11, 2026.
Vertex AI is Google Cloud's managed model platform. Alongside Google's own Gemini family it resells Anthropic Claude, Meta Llama, Mistral, DeepSeek and Qwen as managed endpoints — all billed per token on the same GCP invoice, with IAM, VPC Service Controls and committed-use discounts applying across the board.
Last updated · synced weekly from the upstream model catalog
Anthropic
$1.25/1M
from Claude 3 Haiku
$0.250/1M
from GPT OSS 20B
Meta
$0.700/1M
from Meta Llama 4 Scout 17B 128e Instruct Maas
Mistral
$0.150/1M
from Mistral Nemo Latest
DeepSeek
$1.68/1M
from DeepSeek Ai DeepSeek V3.2 Maas
Qwen
$1.00/1M
from Qwen Qwen3 235B A22b Instruct 2507 Maas
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 104 priced models, cheapest first.
| Model | Vendor | Input / 1M | Output / 1M | Batch out / 1M | Context |
|---|---|---|---|---|---|
| Mistral Nemo Latest | Mistral | $0.150 | $0.150 | — | 128K |
| GPT OSS 20B ReasoningTools | $0.070 | $0.250 | — | 131K | |
| Gemini 2.0 Flash Lite | $0.075 | $0.300 | — | 1.0M | |
| Gemini 2.0 Flash Lite 001 | $0.075 | $0.300 | — | 1.0M | |
| GPT OSS 120B ReasoningTools | $0.090 | $0.360 | — | 131K | |
| Gemini 2.0 Flash | $0.100 | $0.400 | — | 1.0M | |
| Gemini 2.5 Flash Lite Preview 06.17 | $0.100 | $0.400 | — | 1.0M | |
| Gemini 2.5 Flash Lite Preview 09 2025 | $0.100 | $0.400 | — | 1.0M | |
| Gemini 2.5 Flash-Lite VisionReasoningToolsCache | $0.100 | $0.400 | — | 1.0M | |
| Xai Grok 4.1 Fast Non Reasoning | $0.200 | $0.500 | — | 2.0M | |
| Xai Grok 4.1 Fast Reasoning | $0.200 | $0.500 | — | 2.0M | |
| Codestral 2405 | Mistral | $0.200 | $0.600 | — | 128K |
| Codestral 2501 | Mistral | $0.200 | $0.600 | — | 128K |
| Codestral Latest | Mistral | $0.200 | $0.600 | — | 128K |
| Gemini 2.0 Flash 001 | $0.150 | $0.600 | — | 1.0M | |
| Meta Llama 4 Scout 17B 128e Instruct Maas | Meta | $0.250 | $0.700 | — | 10.0M |
| Meta Llama 4 Scout 17B 16e Instruct Maas | Meta | $0.250 | $0.700 | — | 10.0M |
| Llama 3.3 70B Instruct Tools | $0.720 | $0.720 | — | 128K | |
| Qwen3 235B A22B Instruct ReasoningTools | $0.220 | $0.880 | — | 262K | |
| Codestral 2 | Mistral | $0.300 | $0.900 | — | 128K |
| Codestral 2 001 | Mistral | $0.300 | $0.900 | — | 128K |
| Mistralai Codestral 2 | Mistral | $0.300 | $0.900 | — | 128K |
| Mistralai Codestral 2 001 | Mistral | $0.300 | $0.900 | — | 128K |
| Qwen Qwen3 235B A22b Instruct 2507 Maas | Qwen | $0.250 | $1.00 | — | 262K |
| Llama 4 Maverick 17B 128E Instruct VisionTools | $0.350 | $1.15 | — | 524K | |
| Meta Llama 4 Maverick 17B 128e Instruct Maas | Meta | $0.350 | $1.15 | — | 1.0M |
| Meta Llama 4 Maverick 17B 16e Instruct Maas | Meta | $0.350 | $1.15 | — | 1.0M |
| Qwen Qwen3 Next 80B A3b Instruct Maas | Qwen | $0.150 | $1.20 | — | 262K |
| Qwen Qwen3 Next 80B A3b Thinking Maas | Qwen | $0.150 | $1.20 | — | 262K |
| Claude 3 Haiku | Anthropic | $0.250 | $1.25 | — | 200K |
| Gemini 3.1 Flash Lite VisionReasoningToolsCache | $0.250 | $1.50 | $0.750−50% | 1.0M | |
| Gemini 3.1 Flash Lite Preview VisionReasoningToolsCache | $0.250 | $1.50 | — | 1.0M | |
| Gemini Flash-Lite Latest VisionReasoningToolsCache | $0.250 | $1.50 | — | 1.0M | |
| DeepSeek Ai DeepSeek V3.2 Maas | DeepSeek | $0.560 | $1.68 | $0.840−50% | 164K |
| DeepSeek V3.2 ReasoningToolsCache | $0.560 | $1.68 | — | 164K | |
| DeepSeek V3.1 ReasoningTools | $0.600 | $1.70 | — | 164K | |
| Mistral Medium 3 | Mistral | $0.400 | $2.00 | — | 128K |
| Mistral Medium 3 001 | Mistral | $0.400 | $2.00 | — | 128K |
| Mistralai Mistral Medium 3 | Mistral | $0.400 | $2.00 | — | 128K |
| Mistralai Mistral Medium 3 001 | Mistral | $0.400 | $2.00 | — | 128K |
| GLM-4.7 ReasoningTools | $0.600 | $2.20 | — | 200K | |
| Gemini 2.5 Flash VisionReasoningToolsCache | $0.300 | $2.50 | — | 1.0M | |
| Gemini 2.5 Flash Preview 09 2025 | $0.300 | $2.50 | — | 1.0M | |
| Gemini 3.5 Flash Lite VisionReasoningToolsCache | $0.300 | $2.50 | $1.25−50% | 1.0M | |
| Gemini Robotics Er 1.5 Preview | $0.300 | $2.50 | — | 1.0M | |
| Kimi K2 Thinking ReasoningTools | $0.600 | $2.50 | — | 262K | |
| Gemini 3 Flash Preview VisionReasoningToolsCache | $0.500 | $3.00 | — | 1.0M | |
| Mistral Nemo 2407 | Mistral | $3.00 | $3.00 | — | 128K |
| Mistral Small 2503 | Mistral | $1.00 | $3.00 | — | 128K |
| Mistral Small 2503 001 | Mistral | $1.00 | $3.00 | — | 32K |
| GLM-5 ReasoningToolsCache | $1.00 | $3.20 | — | 203K | |
| Qwen Qwen3 Coder 480B A35b Instruct Maas | Qwen | $1.00 | $4.00 | — | 262K |
| Claude 3.5 Haiku | Anthropic | $1.00 | $5.00 | — | 200K |
| Claude Haiku 4.5 VisionReasoningToolsCache | Anthropic | $1.00 | $5.00 | — | 200K |
| Claude Haiku 4.5 VisionReasoningToolsCache | $1.00 | $5.00 | — | 200K | |
| DeepSeek Ai DeepSeek R1 0528 Maas | DeepSeek | $1.35 | $5.40 | — | 65K |
| DeepSeek Ai DeepSeek V3.1 Maas | DeepSeek | $1.35 | $5.40 | — | 164K |
| Mistral Large 2407 | Mistral | $2.00 | $6.00 | — | 128K |
| Mistral Large 2411 | Mistral | $2.00 | $6.00 | — | 128K |
| Mistral Large 2411 001 | Mistral | $2.00 | $6.00 | — | 128K |
| Mistral Large Latest | Mistral | $2.00 | $6.00 | — | 128K |
| Xai Grok 4.20 Non Reasoning | $2.00 | $6.00 | — | 2.0M | |
| Xai Grok 4.20 Reasoning | $2.00 | $6.00 | — | 2.0M | |
| Gemini 3.6 Flash VisionReasoningToolsCache | $1.50 | $7.50 | $3.75−50% | 1.0M | |
| Gemini 3.5 Flash VisionReasoningToolsCache | $1.50 | $9.00 | — | 1.0M | |
| Gemini Flash Latest VisionReasoningToolsCache | $1.50 | $9.00 | — | 1.0M | |
| Gemini Omni Flash Preview | $1.50 | $9.00 | — | 1.0M | |
| Claude Sonnet 5 VisionReasoningToolsCache | Anthropic | $2.00 | $10.00 | — | 1.0M |
| Claude Sonnet 5 VisionReasoningToolsCache | $2.00 | $10.00 | — | 1.0M | |
| Gemini 2.5 Computer Use Preview 10 2025 | $1.25 | $10.00 | — | 128K | |
| Gemini 2.5 Pro VisionReasoningToolsCache | $1.25 | $10.00 | — | 1.0M | |
| Gemini 2.5 Pro Preview Tts | $1.25 | $10.00 | — | 1.0M | |
| Gemini 3 Pro Preview | $2.00 | $12.00 | $6.00−50% | 1.0M | |
| Gemini 3.1 Pro Preview VisionReasoningToolsCache | $2.00 | $12.00 | $6.00−50% | 1.0M | |
| Gemini 3.1 Pro Preview Custom Tools VisionReasoningToolsCache | $2.00 | $12.00 | $6.00−50% | 1.0M | |
| Claude 3 Sonnet | Anthropic | $3.00 | $15.00 | — | 200K |
| Claude 3.5 Sonnet | Anthropic | $3.00 | $15.00 | — | 200K |
| Claude 3.7 Sonnet (2025-02-19) | Anthropic | $3.00 | $15.00 | — | 200K |
| Claude Sonnet 4 VisionReasoningToolsCache | Anthropic | $3.00 | $15.00 | — | 200K |
| Nano Banana Pro VisionReasoning | $2.00 | $120.00 | — | 66K |
A dash in the batch column means our catalog carries no published batch rate for that model — not that the model has no batch tier. Batch coverage in the upstream pricing data is uneven.
Frequently asked questions
How much do Anthropic Claude models cost on Google Vertex AI?
Claude models on Vertex AI start at $1.25 per 1 million output tokens across 19 models, with the current flagship Claude Opus 5 at $25.00. Claude 3 Haiku is the cheapest. Input tokens cost less than output on every model.
Is Claude cheaper on Vertex AI or through the Anthropic API?
Per-token rates are effectively the same on both — Google resells Anthropic models at parity. What differs is everything around the tokens: Vertex bills through your Google Cloud account with IAM, VPC Service Controls, regional pinning and committed-use discounts, while the Anthropic API bills directly and typically ships new models first.
How much does Gemini cost on Vertex AI?
Gemini models on Vertex AI start at $0.250 per 1 million output tokens across 54 models, with the current flagship Nano Banana Pro at $120.00. The long-context models bill at a higher rate above their tier threshold, so a prompt that crosses the boundary costs more per token than the headline price.
Which models are available on Google Vertex AI?
Vertex AI hosts models from Anthropic, Google, Meta, Mistral, DeepSeek, Qwen — 104 priced models in total. Google's own Gemini family sits alongside third-party models offered as managed endpoints, all billed per token under the same Google Cloud invoice.
What is the largest context window available on Vertex AI?
Meta Llama 4 Scout 17B 128e Instruct Maas (Meta) has the largest context window on Vertex AI at 10.0M input tokens. Long context is billed the same per-token as short context on most models, but a few tier the rate above a threshold.
How do I estimate my Vertex AI bill?
Multiply input tokens by the input rate and output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Vertex also charges separately for grounding, embeddings and provisioned throughput, which are not part of the per-token rates listed here.
Head-to-head comparisons
Vertex AI pricing by vendor
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator