All articles

LLM Economics

Token math, model selection, and the true unit cost of every generation.

Moonshot AI's 300 Billion Daily Tokens: What It Signals About the Model Market
LLM Economics
Sep 15, 20263 min read

Moonshot AI's 300 Billion Daily Tokens: What It Signals About the Model Market

A Chinese lab targeting $2B in revenue is notable, but the more useful number for founders is the sheer token throughput reshaping provider economics.

What Nvidia's Vera Rubin Platform Means for Your Token Costs
LLM Economics
Sep 15, 20264 min read

What Nvidia's Vera Rubin Platform Means for Your Token Costs

Better performance-per-watt on the hardware side eventually shows up as lower token prices, but the timeline and who captures the savings first is worth understanding.

The Real Total Cost of an LLM Gateway: Why the License Fee Is Never the Whole Story
LLM Economics
Sep 15, 20263 min read

The Real Total Cost of an LLM Gateway: Why the License Fee Is Never the Whole Story

A vendor comparison of LiteLLM's enterprise pricing against a managed alternative shows why a $250/month license number hides most of the actual 3-year cost.

What Prompt Caching Actually Saves: The Token Math Behind LLM Gateways
LLM Economics
Sep 15, 20264 min read

What Prompt Caching Actually Saves: The Token Math Behind LLM Gateways

Once caching is on, output tokens dominate your bill, so the input-token discounts everyone advertises matter less than you'd think.

Does AI Pricing Really Need to Fall 90%? What the Token-Cost Math Actually Shows
LLM Economics
Sep 15, 20265 min read

Does AI Pricing Really Need to Fall 90%? What the Token-Cost Math Actually Shows

No, prices don't need to crash 90% overnight, but the trend underneath the headline is real, and it changes how you should price any AI feature you ship this year.

If AI Coding Is Not Winner-Take-All, Gross Margin Decides Who Survives
LLM Economics
Sep 9, 20265 min read

If AI Coding Is Not Winner-Take-All, Gross Margin Decides Who Survives

Cognition raised at a $48B valuation, which TechCrunch reads as investors betting that AI coding is not a winner-take-all market, and in a market with many survivors the winners are usually decided by unit economics rather than by being first.

A 33x Price Gap for the Same Model Is Not a Market. It Is a Tax on Not Checking.
LLM Economics
Sep 9, 20265 min read

A 33x Price Gap for the Same Model Is Not a Market. It Is a Tax on Not Checking.

A developer who tracks LLM API list prices daily reports a 33x cost gap for the identical model depending on which provider you call, which means most of what teams call their AI cost is really a procurement decision they never made.

The Same Model, 14 Providers: Why Your LLM Bill Depends on Where You Run It
LLM Economics
Sep 9, 20266 min read

The Same Model, 14 Providers: Why Your LLM Bill Depends on Where You Run It

Choosing a model sets your capability ceiling, but choosing the provider that serves it sets your actual bill, because price, throughput and prompt caching all differ between endpoints running the identical weights.

AI Agent Cost Per Hour: Why the $6 Headline Is a Floor, Not a Bill
LLM Economics
Sep 6, 20265 min read

AI Agent Cost Per Hour: Why the $6 Headline Is a Floor, Not a Bill

An AI agent quoted at under $6 an hour is priced on output tokens at a single stream, so your real bill depends on input volume, how many agents you run in parallel, and how often work has to be redone.

Cost Per Token vs Cost Per Task: Why LLM Pricing Comparisons Keep Misleading Founders
LLM Economics
Sep 5, 20266 min read

Cost Per Token vs Cost Per Task: Why LLM Pricing Comparisons Keep Misleading Founders

A model priced 2.5x higher per token can still produce a smaller bill, because what you actually pay for is tokens consumed per task, not the rate on the pricing page.

Compute Capacity Belongs in Your AI Cost Model, Not Just Your News Feed
LLM Economics
Sep 4, 20266 min read

Compute Capacity Belongs in Your AI Cost Model, Not Just Your News Feed

When a provider adds capacity and raises your rate limits, your effective cost per user can fall even though the list price per token has not moved at all.

A Cheaper Model Does Not Lower Your AI Bill: What Fable 5.1 Actually Changes
LLM Economics
Sep 2, 20266 min read

A Cheaper Model Does Not Lower Your AI Bill: What Fable 5.1 Actually Changes

Anthropic's Fable 5.1 is built to reduce token cost, but your monthly AI bill is set by cost per completed task, not by the number on the model card.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.