All articles
LLM Economics

Does AI Pricing Really Need to Fall 90%? What the Token-Cost Math Actually Shows

No, prices don't need to crash 90% overnight, but the trend underneath the headline is real, and it changes how you should price any AI feature you ship this year.

Sep 15, 2026 · 5 min read
Does AI Pricing Really Need to Fall 90%? What the Token-Cost Math Actually Shows

Key takeaways

  • Frontier model prices have already fallen roughly 10x to 100x on a per-token basis since 2023, depending on the model tier.
  • The "90% must fall further" argument confuses provider list prices with your blended cost per user, which depends on your own token usage, not the sticker price.
  • Margin pressure is real for wrapper products with thin differentiation, not for products where the AI is a small fraction of total value delivered.
  • The safer move is building your pricing model around your own measured token costs per user or per request, not around vendor rate cards.
  • Recompute your unit economics every time a provider changes pricing, not once a year.

Why do people expect AI prices to fall 90%?

A widely shared debate thread argued that AI product pricing has to drop by roughly 90% because the underlying token costs paid by providers are dropping even faster than the prices charged to end customers. The logic: if compute costs keep falling due to better chips and more efficient models, but SaaS-style AI products keep charging subscription prices set when compute was expensive, someone is capturing an unsustainable margin, and competition will eventually force it out.

There's a real pattern behind the claim. Say, for example, a provider's flagship model cost $15 per million input tokens in 2023 and a comparable-capability model costs $2 per million tokens today, that is roughly an 85% drop in raw compute cost for the same job. Layer in prompt caching, smaller distilled models good enough for most tasks, and price wars between providers, and the cost of running the same feature can fall dramatically within 12-18 months.

What's actually driving the cost curve down?

Three forces compound: model efficiency (smaller models matching larger ones on narrower tasks), infrastructure gains (newer chips delivering more tokens per watt), and competitive pricing between providers racing for developer market share. None of these forces are slowing down, so the direction of the trend is not in question. What's in question is whether that maps directly onto a mandate to cut your own product's price by 90%.

Where does the 90% claim break down?

The claim treats "token cost" and "AI product price" as the same thing. They aren't. Your product's price should reflect the value delivered, your margin target, and your actual blended cost per active user, which depends on usage patterns you control (context length, caching, model tier chosen per task) far more than on the provider's list price alone.

For example, say a support-automation tool spends an average of 8,000 tokens per resolved ticket. If the underlying model cost falls from $3 to $0.50 per million tokens, the true compute cost per ticket falls from roughly 2.4 cents to 0.4 cents. That's real margin recovered, not a mandate to cut the customer-facing price by 90%, it's an opportunity to either expand margin or reinvest it into more generous usage limits.

The 90% framing also ignores products where AI is a thin layer over substantial non-AI value: workflow, data, integrations, trust. Those products can absorb falling compute costs as pure margin expansion rather than pass-through pricing cuts.

How should founders price AI features today?

Three practical moves: first, model your own token cost per user or per request using your actual usage data, not vendor marketing numbers. Second, re-run that model every time a provider changes pricing or you swap models, since a 4x cost swing between model tiers is common. Third, separate "cost of the AI" from "value of the outcome" in your own pricing conversations internally, so a falling compute bill becomes a margin lever you control rather than an external price war you're forced into.

Founders who wait for perfect cost clarity before adjusting pricing tend to either overcharge and lose deals to leaner competitors, or undercharge and quietly bleed margin for quarters before noticing. Modeling your blended cost per tier before you ship a price change closes that gap.

The takeaway

AI compute costs are falling fast and will keep falling, but that's a margin opportunity to manage deliberately, not an automatic mandate to slash prices 90%. Calcaas lets you model your own token cost per user or per request against the latest provider rates, so pricing decisions are based on your numbers, not headlines.

Frequently asked questions

Will AI subscription prices really fall 90%?

Not uniformly. Token costs for comparable model capability have fallen sharply, often 80-90% over 12-24 months for a given performance tier, but end-customer pricing depends on margin strategy, differentiation, and how much of the product's value is AI versus everything else around it.

How do I know if my AI feature is overpriced relative to its cost?

Calculate your blended token cost per active user or per completed task using your actual usage logs, then compare that to the price your customer pays for that feature. If the ratio has shifted significantly since you last checked, it's time to revisit the price or the margin target.

Should I pass falling token costs on to customers as lower prices?

Not automatically. Falling costs can fund better limits, added features, or higher margin, all of which serve the business more sustainably than a reflexive price cut that's hard to reverse later.

What's the fastest way to model this without building a spreadsheet from scratch?

A pricing calculator that already has current per-token rates across providers lets you plug in your usage pattern and see blended cost per user immediately, rather than manually tracking rate cards that change monthly. (Note: place this JSON-LD inside a <script type="application/ld+json"> tag in the page head.)

ShareXLinkedInFacebook

More from the blog

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.