Google Vertex AI Pricing
Google Vertex AI resells 102 large language models from 6 vendors, billed per million tokens through Google Cloud. Mistral Nemo Latest (Mistral) is the cheapest at $0.150 per 1M output tokens. Anthropic Claude models on Vertex run $1.25–$25.00 per 1M output. Prices are USD, current as of July 20, 2026.
Vertex AI is Google Cloud's managed model platform. Alongside Google's own Gemini family it resells Anthropic Claude, Meta Llama, Mistral, DeepSeek and Qwen as managed endpoints — all billed per token on the same GCP invoice, with IAM, VPC Service Controls and committed-use discounts applying across the board.
Last updated · synced weekly from the upstream model catalog
Anthropic
$1.25/1M
from Claude 3 Haiku
$0.250/1M
from GPT OSS 20B
Meta
$0.700/1M
from Meta Llama 4 Scout 17B 128e Instruct Maas
Mistral
$0.150/1M
from Mistral Nemo Latest
DeepSeek
$1.68/1M
from DeepSeek Ai DeepSeek V3.2 Maas
Qwen
$1.00/1M
from Qwen Qwen3 235B A22b Instruct 2507 Maas
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 102 priced models, cheapest first.
| Model | Vendor | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| Mistral Nemo Latest | Mistral | $0.150 | $0.150 | 128K |
| GPT OSS 20B ReasoningTools | $0.070 | $0.250 | 131K | |
| Gemini 2.0 Flash Lite | $0.075 | $0.300 | 1.0M | |
| Gemini 2.0 Flash Lite 001 | $0.075 | $0.300 | 1.0M | |
| GPT OSS 120B ReasoningTools | $0.090 | $0.360 | 131K | |
| Gemini 2.0 Flash | $0.100 | $0.400 | 1.0M | |
| Gemini 2.5 Flash Lite Preview 06.17 | $0.100 | $0.400 | 1.0M | |
| Gemini 2.5 Flash Lite Preview 09 2025 | $0.100 | $0.400 | 1.0M | |
| Gemini 2.5 Flash-Lite VisionReasoningToolsCache | $0.100 | $0.400 | 1.0M | |
| Xai Grok 4.1 Fast Non Reasoning | $0.200 | $0.500 | 2.0M | |
| Xai Grok 4.1 Fast Reasoning | $0.200 | $0.500 | 2.0M | |
| Codestral 2405 | Mistral | $0.200 | $0.600 | 128K |
| Codestral 2501 | Mistral | $0.200 | $0.600 | 128K |
| Codestral Latest | Mistral | $0.200 | $0.600 | 128K |
| Gemini 2.0 Flash 001 | $0.150 | $0.600 | 1.0M | |
| Meta Llama 4 Scout 17B 128e Instruct Maas | Meta | $0.250 | $0.700 | 10.0M |
| Meta Llama 4 Scout 17B 16e Instruct Maas | Meta | $0.250 | $0.700 | 10.0M |
| Llama 3.3 70B Instruct Tools | $0.720 | $0.720 | 128K | |
| Qwen3 235B A22B Instruct ReasoningTools | $0.220 | $0.880 | 262K | |
| Codestral 2 | Mistral | $0.300 | $0.900 | 128K |
| Codestral 2 001 | Mistral | $0.300 | $0.900 | 128K |
| Mistralai Codestral 2 | Mistral | $0.300 | $0.900 | 128K |
| Mistralai Codestral 2 001 | Mistral | $0.300 | $0.900 | 128K |
| Qwen Qwen3 235B A22b Instruct 2507 Maas | Qwen | $0.250 | $1.00 | 262K |
| Llama 4 Maverick 17B 128E Instruct VisionTools | $0.350 | $1.15 | 524K | |
| Meta Llama 4 Maverick 17B 128e Instruct Maas | Meta | $0.350 | $1.15 | 1.0M |
| Meta Llama 4 Maverick 17B 16e Instruct Maas | Meta | $0.350 | $1.15 | 1.0M |
| Qwen Qwen3 Next 80B A3b Instruct Maas | Qwen | $0.150 | $1.20 | 262K |
| Qwen Qwen3 Next 80B A3b Thinking Maas | Qwen | $0.150 | $1.20 | 262K |
| Claude 3 Haiku | Anthropic | $0.250 | $1.25 | 200K |
| Gemini 3.1 Flash Lite VisionReasoningToolsCache | $0.250 | $1.50 | 1.0M | |
| Gemini 3.1 Flash Lite Preview VisionReasoningToolsCache | $0.250 | $1.50 | 1.0M | |
| Gemini Flash-Lite Latest VisionReasoningToolsCache | $0.250 | $1.50 | 1.0M | |
| DeepSeek Ai DeepSeek V3.2 Maas | DeepSeek | $0.560 | $1.68 | 164K |
| DeepSeek V3.2 ReasoningToolsCache | $0.560 | $1.68 | 164K | |
| DeepSeek V3.1 ReasoningTools | $0.600 | $1.70 | 164K | |
| Mistral Medium 3 | Mistral | $0.400 | $2.00 | 128K |
| Mistral Medium 3 001 | Mistral | $0.400 | $2.00 | 128K |
| Mistralai Mistral Medium 3 | Mistral | $0.400 | $2.00 | 128K |
| Mistralai Mistral Medium 3 001 | Mistral | $0.400 | $2.00 | 128K |
| GLM-4.7 ReasoningTools | $0.600 | $2.20 | 200K | |
| Gemini 2.5 Flash VisionReasoningToolsCache | $0.300 | $2.50 | 1.0M | |
| Gemini 2.5 Flash Preview 09 2025 | $0.300 | $2.50 | 1.0M | |
| Gemini Robotics Er 1.5 Preview | $0.300 | $2.50 | 1.0M | |
| Kimi K2 Thinking ReasoningTools | $0.600 | $2.50 | 262K | |
| Gemini 3 Flash Preview VisionReasoningToolsCache | $0.500 | $3.00 | 1.0M | |
| Mistral Nemo 2407 | Mistral | $3.00 | $3.00 | 128K |
| Mistral Small 2503 | Mistral | $1.00 | $3.00 | 128K |
| Mistral Small 2503 001 | Mistral | $1.00 | $3.00 | 32K |
| GLM-5 ReasoningToolsCache | $1.00 | $3.20 | 203K | |
| Claude Haiku 3.5 VisionToolsCache | Anthropic | $0.800 | $4.00 | 200K |
| Qwen Qwen3 Coder 480B A35b Instruct Maas | Qwen | $1.00 | $4.00 | 262K |
| Claude 3.5 Haiku | Anthropic | $1.00 | $5.00 | 200K |
| Claude Haiku 4.5 VisionReasoningToolsCache | Anthropic | $1.00 | $5.00 | 200K |
| DeepSeek Ai DeepSeek R1 0528 Maas | DeepSeek | $1.35 | $5.40 | 65K |
| DeepSeek Ai DeepSeek V3.1 Maas | DeepSeek | $1.35 | $5.40 | 164K |
| Mistral Large 2407 | Mistral | $2.00 | $6.00 | 128K |
| Mistral Large 2411 | Mistral | $2.00 | $6.00 | 128K |
| Mistral Large 2411 001 | Mistral | $2.00 | $6.00 | 128K |
| Mistral Large Latest | Mistral | $2.00 | $6.00 | 128K |
| Xai Grok 4.20 Non Reasoning | $2.00 | $6.00 | 2.0M | |
| Xai Grok 4.20 Reasoning | $2.00 | $6.00 | 2.0M | |
| Gemini 3.5 Flash VisionReasoningToolsCache | $1.50 | $9.00 | 1.0M | |
| Gemini Flash Latest VisionReasoningToolsCache | $1.50 | $9.00 | 1.0M | |
| Gemini Omni Flash Preview | $1.50 | $9.00 | 1.0M | |
| Claude Sonnet 5 VisionReasoningToolsCache | Anthropic | $2.00 | $10.00 | 1.0M |
| Gemini 2.5 Computer Use Preview 10 2025 | $1.25 | $10.00 | 128K | |
| Gemini 2.5 Pro VisionReasoningToolsCache | $1.25 | $10.00 | 1.0M | |
| Gemini 2.5 Pro Preview Tts | $1.25 | $10.00 | 1.0M | |
| Gemini 3 Pro Preview | $2.00 | $12.00 | 1.0M | |
| Gemini 3.1 Pro Preview VisionReasoningToolsCache | $2.00 | $12.00 | 1.0M | |
| Gemini 3.1 Pro Preview Custom Tools VisionReasoningToolsCache | $2.00 | $12.00 | 1.0M | |
| Claude 3 Sonnet | Anthropic | $3.00 | $15.00 | 200K |
| Claude 3.5 Sonnet | Anthropic | $3.00 | $15.00 | 200K |
| Claude 3.7 Sonnet (2025-02-19) | Anthropic | $3.00 | $15.00 | 200K |
| Claude Sonnet 4 VisionReasoningToolsCache | Anthropic | $3.00 | $15.00 | 200K |
| Claude Sonnet 4.5 VisionReasoningToolsCache | Anthropic | $3.00 | $15.00 | 200K |
| Claude Sonnet 4.6 VisionReasoningToolsCache | Anthropic | $3.00 | $15.00 | 1.0M |
| Meta Llama 3.1 405B Instruct Maas | Meta | $5.00 | $16.00 | 128K |
| Nano Banana Pro VisionReasoning | $2.00 | $120.00 | 66K |
Frequently asked questions
How much do Anthropic Claude models cost on Google Vertex AI?
Claude models on Vertex AI range from $1.25 to $25.00 per 1 million output tokens across 19 models. Claude 3 Haiku is the cheapest; Claude Opus 4.8 is the most capable and most expensive. Input tokens cost less than output on every model.
Is Claude cheaper on Vertex AI or through the Anthropic API?
Per-token rates are effectively the same on both — Google resells Anthropic models at parity. What differs is everything around the tokens: Vertex bills through your Google Cloud account with IAM, VPC Service Controls, regional pinning and committed-use discounts, while the Anthropic API bills directly and typically ships new models first.
How much does Gemini cost on Vertex AI?
Gemini models on Vertex AI run $0.250 to $120.00 per 1 million output tokens across 52 models. The long-context models bill at a higher rate above their tier threshold, so a prompt that crosses the boundary costs more per token than the headline price.
Which models are available on Google Vertex AI?
Vertex AI hosts models from Anthropic, Google, Meta, Mistral, DeepSeek, Qwen — 102 priced models in total. Google's own Gemini family sits alongside third-party models offered as managed endpoints, all billed per token under the same Google Cloud invoice.
What is the largest context window available on Vertex AI?
Meta Llama 4 Scout 17B 128e Instruct Maas (Meta) has the largest context window on Vertex AI at 10.0M input tokens. Long context is billed the same per-token as short context on most models, but a few tier the rate above a threshold.
How do I estimate my Vertex AI bill?
Multiply input tokens by the input rate and output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Vertex also charges separately for grounding, embeddings and provisioned throughput, which are not part of the per-token rates listed here.
Head-to-head comparisons
Vertex AI pricing by vendor
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator