Jevons Paradox in AI: Why Cheaper Tokens Mean Bigger Bills
Jevons paradox in AI means that when tokens get cheaper, people use so many more that total spend rises. If price per task halves from $0.10 to $0.05, volume must double just to hold spend flat, and a tripling lifts spend 50 percent.

Last updated: October 2026
Key takeaways
- a16z's State of Markets II argues AI has become an Everything Cycle reshaping capex, debt, power and labor.
- It highlights rising cloud backlogs, resilient GPU rental rates and Jevons Paradox.
- Cheaper AI creates more demand, so unit price falls and total spend can still rise.
- Budget on volume growth, not only on the price list.
- Price your product so success does not eat your margin.
What is Jevons paradox in AI?
Jevons paradox says that making a resource more efficient can increase total consumption of it, because lower cost unlocks new uses. In AI, falling prices per token let teams add longer prompts, more agent steps, more background jobs and more features. a16z's 90-page semiannual report, State of Markets II, cites the idea alongside rising cloud backlogs and resilient GPU rental rates, and argues AI has become a macroeconomic Everything Cycle. Read the report at a16z State of Markets II. Note the published date was reported as October 1, 2026 by secondary coverage and was not verified on the source.
Does cheaper AI really increase total spend?
It does when demand is elastic enough. Say a task costs $0.10 today and your product runs 1M tasks a month, so the bill is $100k. These numbers are illustrative.
| Price per task | Monthly volume | Monthly spend | Change vs baseline | |---|---|---|---| | $0.10 (baseline) | 1M | $100k | 0% | | $0.05 | 1.5M | $75k | -25% | | $0.05 | 2M | $100k | 0% | | $0.05 | 3M | $150k | +50% |
Cutting the price in half only saves money if volume grows less than 2x. If usage triples, you spend more even though every task got cheaper. Our look at LLM prices falling 13x in four months showed the unit-price side of this story. The volume side is what a16z points to.
Why are GPU rental rates holding up?
The report highlights resilient GPU rental rates, which fits the same logic. If cheaper inference creates more demand, capacity stays tight and hardware keeps its value. For founders, that means the cost floor under your provider may be firmer than the falling price list suggests. Memory costs add pressure too, as we covered in memory prices and AI inference costs.
How should you budget when prices keep falling?
Three rules help:
- 1Forecast spend as price times volume, with a growth scenario for volume, not just a price scenario.
- 2Track cost per user and per task monthly, so you see volume creep early.
- 3Tie pricing to usage with credits or tiers, so growth in consumption brings revenue with it. See token economics and falling LLM prices for more on that mindset.
The trap is celebrating a price cut and forgetting to cap the new behaviors it enables, such as unlimited agent loops.
What does Jevons mean for your pricing page?
If your plan is a flat fee and your users respond to cheaper AI by using far more of it, your margin is the thing that gets consumed. Flat unlimited plans transfer Jevons risk to you. Usage-aware plans transfer it to the customer. Compare current provider rates on the LLM pricing hub to see how fast your unit cost is moving.
Takeaway: unit prices fall, total bills rise, so budget on volume and price your plans on usage.
To stress-test your forecast, try the LLM cost calculator with a higher volume scenario.
Frequently asked questions
What is Jevons paradox in AI?
It is the pattern where cheaper AI leads to so much more usage that total spending rises. Lower cost per token unlocks new uses such as agents and longer prompts.
Does cheaper AI mean lower AI spend?
Not necessarily. If the price per task halves, volume must double just to keep spend flat, and any faster growth increases the bill.
Why does a16z talk about Jevons paradox for AI?
Its State of Markets II report points to cheaper AI creating more demand, alongside rising cloud backlogs and resilient GPU rental rates, as part of an AI Everything Cycle.
How do I protect margins from Jevons paradox?
Forecast with a volume growth scenario, track cost per task, and use usage-based pricing or credits so more consumption brings more revenue.
FAQ schema (JSON-LD)
Place this inside a <script type="application/ld+json"> tag in the page head.
More from the blog
LLM EconomicsOpenRouter Pricing: What a Gateway Really Costs
OpenRouter pricing is the underlying model rate plus whatever platform fee the gateway adds, so the real question is net cost. Say a 5 percent fee on $10,000 of monthly tokens adds $500, and routing 15 percent of spend to cheaper models saves $1,500, a net gain of $1,000.
LLM EconomicsMoE Cost per Token: Why Active Parameters Rule
For a mixture-of-experts model, compute per token follows active parameters, while the GPUs you must rent follow total parameters. A model with 40B active out of 400B total can be cheaper per token than a dense 70B at high utilization, and more expensive when traffic is thin.
LLM EconomicsJev AI Model Cost: Is It Cheaper per Task?
Jev is a new class of AI model from a ChatGPT co-inventor that promises cheaper and faster software intelligence, but no price list is confirmed here. Judge it by cost per successful task: a model at $0.04 per task with 70 percent success beats a $0.10 incumbent only if failures escalate cheaply.
Pricing math, in your inbox.
One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.