All articles
Pricing Strategy

Token Volume Is Up 25x and Mid-Tier Models Do 90% of the Job. Your Pricing Has to Move.

Tom's Hardware reports token volume exploding 25-fold while mid-tier models deliver roughly 90% of flagship capability at one-sixth the cost, which means the default choice of flagship-everything is now a margin decision rather than a quality one.

Sep 9, 2026 · 6 min read
Token Volume Is Up 25x and Mid-Tier Models Do 90% of the Job. Your Pricing Has to Move.

Key takeaways

  • Token volume is reported up 25-fold, so per-token cost decisions now compound across far more traffic than when most pricing was set.
  • Mid-tier models are reported at about 90% of flagship capability for roughly one-sixth of the cost.
  • That ratio makes flagship-by-default an explicit purchase of the last 10% of capability at six times the price.
  • Volume growth and unit price decline pull margin in opposite directions. Which one wins depends on your routing, not on the market.
  • Most SaaS pricing was set against a token cost that no longer exists.

What the numbers say

Two figures do the work here. Token consumption is reported to have grown 25-fold, and mid-tier models are reported to deliver around 90% of flagship capability at one-sixth the cost. Read together they describe a pricing reckoning: providers are selling far more units at falling effective prices, and the premium tier's value proposition has narrowed to a thin band of capability.

For anyone building on top of these APIs, that is not industry news, it is a direct input to gross margin.

Why does a 25x volume increase change your pricing?

Because cost per unit only matters in proportion to units consumed. When your product made a handful of calls per session, a difference of a few cents per thousand calls was noise, and picking the best model regardless of price was rational. At 25 times the volume, that same per-unit difference is a line item large enough to move gross margin by points.

The practical consequence is that the model routing decision has migrated out of engineering and into finance, whether or not anyone told finance.

Is 90% of capability good enough?

That is the wrong question, and the useful original point here is that 90% is not a property of the model, it is a property of the task.

A model at 90% of flagship capability is effectively 100% on classification, extraction, routing, summarization of clean input, and most structured transformations. It is meaningfully below 100% on long multi-step reasoning, ambiguous instructions, and tasks where a single wrong answer is expensive to recover from.

So the honest framing is not flagship versus mid-tier. It is: what share of your calls are actually hard? For most products the answer is a small minority, and those are the only calls that justify a six times cost multiple. Routing by task difficulty captures most of the savings while protecting the outcomes that matter.

What does that do to gross margin?

An illustrative example, since your real numbers depend on your token shape. Say a $30 per month plan currently spends $9 per user on inference, for 70% gross margin. If 80% of those calls are routine and move to a model at one-sixth the cost, the inference line drops to roughly $3, and gross margin moves toward 90%.

Nothing about the product changed. No price increase, no churn risk, no new feature. The same work is just being bought at the correct tier. That is the cheapest margin improvement available to most AI products right now, and it is available exactly once per company, so it is worth doing carefully.

Should you cut your price because inference got cheaper?

Not automatically, and this is where a lot of founders get it wrong in both directions.

Falling inference cost is not a reason to cut price if your pricing is anchored to customer value rather than to your COGS. It is a reason to either expand margin, expand usage limits at the same price, or move upmarket into workloads that were previously uneconomic. Cutting price because cost fell hands the entire benefit to customers who were not asking for it.

The one case where price should move is a usage-based model where your price is explicitly a markup on tokens. There, a falling floor with a fixed markup means your customers will eventually notice, and a competitor will tell them.

What to do this quarter

  1. 1Classify your calls into hard and routine. Sample 200 real requests, do not estimate.
  2. 2Run the routine set against a mid-tier model and score it on your own evaluations, not on a public benchmark.
  3. 3Re-model gross margin per user with a routed mix rather than a single model.
  4. 4Decide deliberately whether the recovered margin becomes profit, higher limits, or a new tier. Do not let it drift.

Token volume grew 25x while the cheap option got good enough for most of the work, so the question is no longer which model is best, it is which model each call deserves. You can compare capability against cost per model in the Calcaas model leaderboard before you route anything.

Frequently asked questions

Are mid-tier models good enough to replace flagship models?

For routine work, usually yes. Classification, extraction, routing and summarization of clean input rarely need frontier capability. Long multi-step reasoning, ambiguous instructions and high-cost-of-error tasks are where the remaining capability gap still shows up.

What is model routing and why does it save money?

Model routing sends each request to the cheapest model that can handle it, rather than sending everything to one model. Because most products have a small minority of genuinely hard calls, routing captures most of the available savings without degrading the outcomes that matter.

Should I lower my prices because AI inference got cheaper?

Only if your pricing is explicitly a markup on tokens. If your price is anchored to the value the customer receives, falling inference cost is a margin, limits or packaging decision, not a signal to discount.

How does rising token volume affect SaaS gross margin?

It amplifies whatever your per-unit cost decision was. A small per-call cost difference that was noise at low volume becomes a multi-point swing in gross margin once volume grows by an order of magnitude, which is why old pricing assumptions quietly stop holding.

How do I test a cheaper model without risking quality?

Sample a few hundred real production requests, run them through the cheaper model, and score the outputs against your own task-specific evaluations rather than a public leaderboard. Route only the categories that pass, and keep the expensive model for the rest. Place the JSON-LD block above inside a <script type="application/ld+json"> tag in the page head. The questions and answers must stay identical to the visible FAQ section. Source: Tom's Hardware, https://www.tomshardware.com/tech-industry/artificial-intelligence/frontier-ai-faces-pricing-reckoning-as-token-volume-explodes-25-fold-mid-tier-models-deliver-90-percent-of-flagship-capability-at-one-sixth-the-cost

ShareXLinkedInFacebook

More from the blog

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.