Gemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins
Google's new Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, undercutting 3.5 Flash on price while using roughly 17% fewer output tokens per task, a combination that compounds into a bigger margin gain than the sticker price alone suggests.
Jul 23, 2026 · 5 min read
Key takeaways
Gemini 3.6 Flash: $1.50/1M input tokens, $7.50/1M output tokens, priced below 3.5 Flash while cutting output token usage by about 17% on the Artificial Analysis Index (up to 65% on some coding benchmarks).
Gemini 3.5 Flash-Lite: $0.30/1M input, $2.50/1M output tokens, 350 output tokens/second, built for high-throughput agentic and document-processing workloads.
Gemini 3.5 Flash Cyber: a specialized, security-focused model paired with Google's CodeMender agent, limited to governments and trusted partners, no public price.
The real margin story is price per token times tokens per task: a cheaper model that also uses fewer tokens delivers a compounding, not additive, cost reduction.
Founders building agents should model a blended cost across a routing mix (Flash-Lite for volume, 3.6 Flash for complex steps) rather than a single per-model price.
What do Gemini 3.6 Flash and 3.5 Flash-Lite actually cost?
Google priced 3.6 Flash at $1.50 per 1M input tokens and $7.50 per 1M output tokens, below 3.5 Flash on a straight per-token basis. Flash-Lite comes in far cheaper at $0.30/1M input and $2.50/1M output tokens, aimed at high-volume, latency-sensitive work like agentic search and document processing. Neither number applies to 3.5 Flash Cyber: Google is keeping that model restricted to governments and trusted partners through its CodeMender security agent, so there is no public rate card for it yet.
Why does token efficiency matter more than the sticker price?
Google says 3.6 Flash uses about 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on some coding benchmarks like DeepSWE. That is the detail worth sitting with: when a model is both cheaper per token and needs fewer tokens to finish the same task, the two effects multiply rather than add. Say a task cost X in tokens at the old price and token count; the new cost is not simply X times the price ratio, because the token count itself shrank too. For a founder tracking margins, that means price per million tokens alone is not a reliable proxy for cost. Cost per completed task is.
How should you model this in your own pricing?
Most usage-based AI pricing gets built around the provider's per-token rate card. That undercounts what actually moves your margin: the combination of rate and consumption per unit of value delivered (a completed ticket, a generated report, a resolved support case). When a provider cuts both at once, as Google is doing here, your effective cost per task can fall by more than the headline discount implies. That is exactly the kind of shift worth re-running through a cost simulator before you touch your pricing page.
Should you route different tasks to different Gemini models?
The three-model spread, Flash-Lite, 3.6 Flash, and the specialized Cyber model, is itself a pricing signal. Google is optimizing for different points on the cost-latency-quality curve rather than shipping one model for everything. For teams building agents, that argues for routing logic instead of a single default model: cheap, fast Flash-Lite for high-volume or low-complexity steps, 3.6 Flash for multi-step reasoning and coding work where the token savings are largest, and specialized models only where the task demands it. The number that actually matters for your margin is the blended average cost per task across that routing mix, not any single model's rate card.
What does this mean for AI providers competing on price?
Flash-Lite's $0.30/$2.50 pricing sits at the aggressive end of the "mini" tier most major providers now offer, continuing a trend where entry-level models keep getting cheaper and more capable at the same time. If that trend holds, the cost floor for high-volume agentic workloads keeps dropping, which is good for margins on usage-based products and bad for anyone whose pricing model assumed token costs would stay flat.
Modeling a multi-model routing mix like this, with different rates and different token-efficiency assumptions per step, is exactly the kind of scenario you can simulate in Calcaas before committing to a pricing tier.
Frequently asked questions
How much does Gemini 3.6 Flash cost per million tokens?
Gemini 3.6 Flash costs $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, according to Google's announcement. That is priced below Gemini 3.5 Flash on a per-token basis.
How much does Gemini 3.5 Flash-Lite cost?
Gemini 3.5 Flash-Lite is priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, and runs at roughly 350 output tokens per second according to the Artificial Analysis Index.
Is Gemini 3.6 Flash cheaper than Gemini 3.5 Flash?
Yes. Google says 3.6 Flash is priced lower than 3.5 Flash and also uses about 17% fewer output tokens per task on the Artificial Analysis Index, so the effective cost reduction per completed task is larger than the price cut alone.
Is Gemini 3.5 Flash Cyber publicly available?
No. 3.5 Flash Cyber is shipping as a limited-access pilot through Google's CodeMender security agent, available only to governments and trusted partners, so there is no public pricing for it.
Where can I use Gemini 3.6 Flash and 3.5 Flash-Lite?
Both are available now through the Gemini API via Google AI Studio and Android Studio, in the Gemini Enterprise Agent Platform, and in the Gemini consumer app. Flash-Lite is also rolling out in Google Search. Note: place the JSON-LD above inside a <script type="application/ld+json"> tag in the page <head>.