API pricing

Claude on Vertex AI Pricing

Claude on Vertex AI charges between $1.25 and $25.00 per 1 million output tokens, depending on the model. Claude 3 Haiku is the cheapest at $1.25/1M output and $0.250/1M input; Claude Opus 4.8 is the most expensive at $25.00/1M output. The largest context window is 1.0M tokens (Claude Fable 5). Prices are USD, current as of July 20, 2026.

Google Cloud resells Anthropic Claude models through Vertex AI, billed per million tokens under your GCP account rather than an Anthropic API key. Model-for-model the token rates track Anthropic direct pricing; what differs is billing, quota, region availability and committed-use discounting.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.250/1M

Claude 3 Haiku

Cheapest output

$1.25/1M

Claude 3 Haiku

Longest context

1.0M

Claude Fable 5

Models priced

19

Avg $26.71/1M output

Pricing by model

All prices in USD per 1 million tokens. All 19 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Claude 3 Haiku$0.250$1.25200K
Claude Haiku 3.5
VisionToolsCache
$0.800$4.00200K
Claude 3.5 Haiku$1.00$5.00200K
Claude Haiku 4.5
VisionReasoningToolsCache
$1.00$5.00200K
Claude Sonnet 5
VisionReasoningToolsCache
$2.00$10.001.0M
Claude 3 Sonnet$3.00$15.00200K
Claude 3.5 Sonnet$3.00$15.00200K
Claude 3.7 Sonnet (2025-02-19)$3.00$15.00200K
Claude Sonnet 4
VisionReasoningToolsCache
$3.00$15.00200K
Claude Sonnet 4.5
VisionReasoningToolsCache
$3.00$15.00200K
Claude Sonnet 4.6
VisionReasoningToolsCache
$3.00$15.001.0M
Claude Opus 4.5
VisionReasoningToolsCache
$5.00$25.00200K
Claude Opus 4.6
VisionReasoningToolsCache
$5.00$25.001.0M
Claude Opus 4.7
VisionReasoningToolsCache
$5.00$25.001.0M
Claude Opus 4.8
VisionReasoningToolsCache
$5.00$25.001.0M
Claude Fable 5$10.00$50.001.0M
Claude 3 Opus$15.00$75.00200K
Claude Opus 4
VisionReasoningToolsCache
$15.00$75.00200K
Claude Opus 4.1
VisionReasoningToolsCache
$15.00$75.00200K

Frequently asked questions

How much does Claude on Vertex AI cost per 1M tokens?

Claude on Vertex AI pricing ranges from $1.25 to $25.00 per 1 million output tokens across 19 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Claude on Vertex AI model?

Claude 3 Haiku is the cheapest Claude on Vertex AI model on both axes — $0.250 per 1M input tokens and $1.25 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Claude on Vertex AI context window?

Claude Fable 5 has the largest context window in the Claude on Vertex AI catalog at 1.0M input tokens, with up to 128K output tokens per response.

Does Claude on Vertex AI support prompt caching?

Yes. Claude Haiku 3.5 reads cached input at $0.080 per 1M tokens, against $0.800 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Claude on Vertex AI models support reasoning or vision?

The Claude on Vertex AI catalog includes 11 reasoning models, 12 vision models, 12 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Claude on Vertex AI bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator