Pricing comparison

OpenAI vs Google Vertex AI Pricing

Google Vertex AI is cheaper than OpenAI at the entry level: Gemma 3n 4B costs $0.040 per 1M output tokens against $0.140 for gpt-oss-20b — 3.5× the price. On current flagships, OpenAI's GPT-5.6 costs $30.00/1M output against $30.00/1M for Google Vertex AI's Nano Banana 2 Lite. Prices are USD, current as of August 11, 2026.

Per-million-token pricing for OpenAI and Google Vertex AI, with side-by-side flagship models, cheapest tiers, and context windows. Pricing data syncs weekly from a continuously-updated model catalog — last updated August 11, 2026.

Who wins on what

Cheapest input tokens

$0.02/1M

Google Vertex AI

Gemma 3 4B — $0.02/1M input

Cheapest output tokens

$0.04/1M

Google Vertex AI

Gemma 3n 4B — $0.04/1M output

Longest context window

2.0M

OpenAI

gpt-5.4 (>272K context length) — 2.0M input tokens

Lowest average output cost

$9.45/1M

Google Vertex AI

Provider-wide average across 59 models

Largest model catalog

129 models

OpenAI

More options to match cost vs capability

Cheapest cached input

$0.02/1M

OpenAI

GPT-5.6 Luna — $0.02/1M cached read

Most reasoning models

18 models

Google Vertex AI

Models with dedicated reasoning / thinking support

Most vision models

18 models

Google Vertex AI

Models that accept image input

Side-by-side

129 models

OpenAI

Full OpenAI pricing →

Cheapest input

$0.030

gpt-oss-20b

Cheapest output

$0.140

gpt-oss-20b

Longest context

2.0M

gpt-5.4 (>272K context length)

Avg output / 1M

$29.47

Across catalog

Cheapest cached input

$0.020

GPT-5.6 Luna

Batch discount

50% off output

35 models with published batch rates

ModelIn/1MOut/1MCtx
GPT-5.6
VisionReasoningToolsCache
Batch −50%
$5.00$30.001.1M
GPT-5.6 Sol
VisionReasoningToolsCache
Batch −50%
$5.00$30.001.1M
GPT-5.6 Terra
VisionReasoningToolsCache
Batch −50%
$2.00$12.001.1M
GPT-5.6 Luna
VisionReasoningToolsCache
Batch −50%
$0.200$1.201.1M
GPT-Realtime-2.1
VisionReasoningToolsCache
$4.00$24.00128K
gpt-oss-20b$0.030$0.140131K
59 models

Google Vertex AI

Full Google Vertex AI pricing →

Cheapest input

$0.017

Gemma 3 4B

Cheapest output

$0.040

Gemma 3n 4B

Longest context

2.0M

Gemini 3/3.1 (> 200k context)

Avg output / 1M

$9.45

Across catalog

Cheapest cached input

$0.025

Gemini 3.1 Flash Lite

Batch discount

50% off output

6 models with published batch rates

ModelIn/1MOut/1MCtx
Gemini 3.6 Flash
VisionReasoningToolsCache
Batch −50%
$1.50$7.501.0M
Gemini 3.5 Flash Lite
VisionReasoningToolsCache
Batch −50%
$0.300$2.501.0M
Nano Banana 2 Lite
VisionReasoningTools
$0.250$30.0066K
Gemini 3.5 Live Translate Preview$3.50$21.0016K
Gemini 3.5 Flash
VisionReasoningToolsCache
$1.50$9.001.0M
Gemma 3n 4B$0.020$0.04033K

All prices in USD per 1 million tokens. Showing top 6 models per provider, sorted by output cost.

Batch badges mark models with a published asynchronous batch rate. A model without one is billed at its standard rate — our catalog does not carry batch rates for every model, so an absent badge is not evidence that no batch tier exists.

Frequently asked questions

Is OpenAI or Google Vertex AI cheaper?

Google Vertex AI has the cheaper entry point at $0.04/1M output (Gemma 3n 4B — $0.04/1M output). Provider-wide, OpenAI averages $29.47/1M output against $9.45/1M for Google Vertex AI. Which is cheaper for you depends on which model tier your workload actually needs.

How much do OpenAI and Google Vertex AI cost per 1M tokens?

OpenAI starts at $0.140 per 1M output tokens (gpt-oss-20b), with its current flagship GPT-5.6 at $30.00. Google Vertex AI starts at $0.040 (Gemma 3n 4B), with Nano Banana 2 Lite at $30.00. Input tokens cost less than output on both.

Which has the larger context window, OpenAI or Google Vertex AI?

OpenAI — gpt-5.4 (>272K context length) — 2.0M input tokens. For comparison, OpenAI's largest is 2.0M tokens (gpt-5.4 (>272K context length)) and Google Vertex AI's is 2.0M tokens (Gemini 3/3.1 (> 200k context)).

Which has more reasoning models, OpenAI or Google Vertex AI?

OpenAI lists 12 reasoning models and Google Vertex AI lists 18. Reasoning models bill their internal thinking as output tokens, so a reasoning call costs several times a standard completion of the same visible length — compare them on total tokens billed, not headline rate.

Do OpenAI and Google Vertex AI offer batch discounts?

Both publish discounted rates for asynchronous batch jobs: OpenAI on 35 models in this catalog and Google Vertex AI on 6, both at 50% off output tokens. Batch trades real-time responses for a lower rate, so it fits backfills, evaluations and bulk enrichment rather than anything user-facing.

Should I switch from OpenAI to Google Vertex AI to save money?

Only if the cheaper model still meets your quality bar. Token price is one input; the ones that decide your bill are prompt size, response length, retries and how much conversation history you resend each turn. Model the switch against your real traffic before committing. Prices here are current as of August 11, 2026.

Related comparisons

See the full AI model leaderboard — every model ranked by intelligence, real-world usage, and value.

Run the numbers for your workload

Calcaas multiplies per-token costs by your real usage patterns — inputs, outputs, retries, and conversation history — across both providers in one model.