OpenAI vs Google Vertex AI Pricing
Google Vertex AI is cheaper than OpenAI at the entry level: Gemma 3n 4B costs $0.040 per 1M output tokens against $0.140 for gpt-oss-20b — 3.5× the price. On current flagships, OpenAI's GPT-6 Astra costs $50.00/1M output against $3.75/1M for Google Vertex AI's Gemini 3.8 Flash. Prices are USD, current as of September 21, 2026.
Per-million-token pricing for OpenAI and Google Vertex AI, with side-by-side flagship models, cheapest tiers, and context windows. Pricing data syncs weekly from a continuously-updated model catalog — last updated September 21, 2026.
Who wins on what
Cheapest input tokens
$0.02/1MGoogle Vertex AI
Gemma 3 4B — $0.02/1M input
Cheapest output tokens
$0.04/1MGoogle Vertex AI
Gemma 3n 4B — $0.04/1M output
Longest context window
2.0MOpenAI
gpt-5.4 (>272K context length) — 2.0M input tokens
Lowest average output cost
$9.61/1MGoogle Vertex AI
Provider-wide average across 65 models
Largest model catalog
137 modelsOpenAI
More options to match cost vs capability
Cheapest cached input
$0.02/1MOpenAI
GPT-5.6 Luna — $0.02/1M cached read
Most reasoning models
19 modelsGoogle Vertex AI
Models with dedicated reasoning / thinking support
Most vision models
19 modelsGoogle Vertex AI
Models that accept image input
Side-by-side
OpenAI
Full OpenAI pricing →Cheapest input
$0.030
gpt-oss-20b
Cheapest output
$0.140
gpt-oss-20b
Longest context
2.0M
gpt-5.4 (>272K context length)
Avg output / 1M
$30.24
Across catalog
Cheapest cached input
$0.020
GPT-5.6 Luna
Batch discount
17–50% off output
58 models with published batch rates
| Model | In/1M | Out/1M | Ctx |
|---|---|---|---|
| GPT-6 Astra VisionReasoningToolsCache Batch −50% | $10.00 | $50.00 | 1.1M |
| GPT-5.6 VisionReasoningToolsCache Batch −50% | $4.00 | $20.00 | 1.1M |
| GPT-5.6 Sol VisionReasoningToolsCache Batch −50% | $4.00 | $20.00 | 1.1M |
| GPT-5.6 Terra VisionReasoningToolsCache Batch −50% | $2.00 | $12.00 | 1.1M |
| GPT-5.6 Luna VisionReasoningToolsCache Batch −50% | $0.200 | $1.20 | 1.1M |
| gpt-oss-20b | $0.030 | $0.140 | 131K |
Google Vertex AI
Full Google Vertex AI pricing →Cheapest input
$0.017
Gemma 3 4B
Cheapest output
$0.040
Gemma 3n 4B
Longest context
2.0M
Gemini 3/3.1 (> 200k context)
Avg output / 1M
$9.61
Across catalog
Cheapest cached input
$0.025
Gemini 3.1 Flash Lite
Batch discount
50% off output
18 models with published batch rates
| Model | In/1M | Out/1M | Ctx |
|---|---|---|---|
| Gemini 3.8 Flash VisionReasoningToolsCache Batch −50% | $0.750 | $3.75 | 1.0M |
| Gemini 3.7 Flash VisionReasoningToolsCache Batch −50% | $0.750 | $3.75 | 1.0M |
| Gemini Flash Latest VisionReasoningToolsCache Batch −50% | $0.750 | $3.75 | 1.0M |
| Gemini 3.6 Flash VisionReasoningToolsCache Batch −50% | $0.750 | $3.75 | 1.0M |
| Gemini 3.5 Flash Lite VisionReasoningToolsCache Batch −50% | $0.300 | $2.50 | 1.0M |
| Gemma 3n 4B | $0.020 | $0.040 | 33K |
All prices in USD per 1 million tokens. Showing top 6 models per provider, sorted by output cost.
Batch badges mark models with a published asynchronous batch rate. A model without one is billed at its standard rate — our catalog does not carry batch rates for every model, so an absent badge is not evidence that no batch tier exists.
Frequently asked questions
Is OpenAI or Google Vertex AI cheaper?
Google Vertex AI has the cheaper entry point at $0.04/1M output (Gemma 3n 4B — $0.04/1M output). Provider-wide, OpenAI averages $30.24/1M output against $9.61/1M for Google Vertex AI. Which is cheaper for you depends on which model tier your workload actually needs.
How much do OpenAI and Google Vertex AI cost per 1M tokens?
OpenAI starts at $0.140 per 1M output tokens (gpt-oss-20b), with its current flagship GPT-6 Astra at $50.00. Google Vertex AI starts at $0.040 (Gemma 3n 4B), with Gemini 3.8 Flash at $3.75. Input tokens cost less than output on both.
Which has the larger context window, OpenAI or Google Vertex AI?
OpenAI — gpt-5.4 (>272K context length) — 2.0M input tokens. For comparison, OpenAI's largest is 2.0M tokens (gpt-5.4 (>272K context length)) and Google Vertex AI's is 2.0M tokens (Gemini 3/3.1 (> 200k context)).
Which has more reasoning models, OpenAI or Google Vertex AI?
OpenAI lists 13 reasoning models and Google Vertex AI lists 19. Reasoning models bill their internal thinking as output tokens, so a reasoning call costs several times a standard completion of the same visible length — compare them on total tokens billed, not headline rate.
Do OpenAI and Google Vertex AI offer batch discounts?
Both publish discounted rates for asynchronous batch jobs: OpenAI on 58 models in this catalog and Google Vertex AI on 18, at 17–50% and 50% off output tokens respectively. Batch trades real-time responses for a lower rate, so it fits backfills, evaluations and bulk enrichment rather than anything user-facing.
Should I switch from OpenAI to Google Vertex AI to save money?
Only if the cheaper model still meets your quality bar. Token price is one input; the ones that decide your bill are prompt size, response length, retries and how much conversation history you resend each turn. Model the switch against your real traffic before committing. Prices here are current as of September 21, 2026.
Related comparisons
Run the numbers for your workload
Calcaas multiplies per-token costs by your real usage patterns — inputs, outputs, retries, and conversation history — across both providers in one model.