LLM Economics

LLM Economics

Token math, model selection, and the true unit cost of every generation.

Page 3 of 5

The $3 Trillion AI Question and the Only Version of It You Can Answer
LLM Economics
Jul 14, 20264 min read

The $3 Trillion AI Question and the Only Version of It You Can Answer

Sequoia's David Cahn now estimates 2026 AI infrastructure spending at $1.5 trillion, implying roughly $3 trillion in revenue to justify it, a question no founder can answer at industry scale but every founder must answer at product scale.

LLM Economics
Jul 14, 20264 min read

Token Economics for AI Budgets: Why Falling LLM Prices Don't Lower Your Bill

The FinOps Foundation argues that tokens are the atomic unit of AI value, and that consumption growth and workload mix, not falling per-token list prices, now determine what organizations actually spend.

The $165K Rewrite: What Bun's 11-Day AI Migration Says About Token Economics
LLM Economics
Jul 14, 20264 min read

The $165K Rewrite: What Bun's 11-Day AI Migration Says About Token Economics

Bun's team compressed a Zig-to-Rust migration estimated at 1-2 years of engineering into 11 days by spending roughly $165K on AI coding agents, one of the clearest public datapoints yet on what large token budgets actually buy.

GPU Cost Per Million Tokens in 2026: What Self-Hosting an LLM Really Costs
LLM Economics
Jul 14, 20264 min read

GPU Cost Per Million Tokens in 2026: What Self-Hosting an LLM Really Costs

Fresh benchmarks across five GPU types put self-hosted LLM inference between $0.16 and $3.58 per million tokens, but your real cost is set by utilization, not by the GPU's hourly rate.

Open Source vs Frontier AI: Why Volume and Spend Tell Opposite Stories
LLM Economics
Jul 9, 20264 min read

Open Source vs Frontier AI: Why Volume and Spend Tell Opposite Stories

Short answer: open-source models are winning token volume while frontier labs keep most of the spend, because the two serve different phases of the same lifecycle, discovery versus production.

Provider Substitution: The AI Cost-Cutting Lever Microsoft Just Pulled
LLM Economics
Jul 8, 20264 min read

Provider Substitution: The AI Cost-Cutting Lever Microsoft Just Pulled

Microsoft is now routing a share of Excel and Word prompts to its own MAI models instead of OpenAI and Anthropic, a direct move to cut AI costs and protect margins.

The Token Apocalypse: Surviving Runaway AI Agent Token Costs
LLM Economics
Jul 7, 20264 min read

The Token Apocalypse: Surviving Runaway AI Agent Token Costs

Short answer: AI agents multiply token consumption by looping, retrying, and chaining calls, so the fastest way to protect margins is to route cheaper work to the right model and model provider switches against your real usage before costs spiral.

The Real Cost of AI: What Google and Amazon's Emissions Spike Signals for Your Token Margins
LLM Economics
Jul 6, 20264 min read

The Real Cost of AI: What Google and Amazon's Emissions Spike Signals for Your Token Margins

The advertised price per token is not the real cost of AI: Google's carbon emissions jumped 25% and Amazon's 16% in a year, a signal that the energy behind every inference call is getting more expensive, not less.

The Inference Inflection: Why AI Margins Now Live in Tokens, Not Training
LLM Economics
Jul 2, 20264 min read

The Inference Inflection: Why AI Margins Now Live in Tokens, Not Training

In short: the cost and margin of an AI product have moved from one-time training to per-request inference, so your unit economics now rise and fall with token costs.

Claude Sonnet 5 Pricing: What $2/$10 per Million Tokens Means for Your Margins
LLM Economics
Jul 1, 20264 min read

Claude Sonnet 5 Pricing: What $2/$10 per Million Tokens Means for Your Margins

Claude Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens (introductory, through August 31, 2026), then $3/$15 - but a new tokenizer means your real cost depends on tokens per task, not the sticker rate.

The Economy of Tokens: Why Faster Inference Doesn't Always Cut Your AI Bill
LLM Economics
Jun 30, 20264 min read

The Economy of Tokens: Why Faster Inference Doesn't Always Cut Your AI Bill

Faster inference frameworks like DeepSeek's DSpark speed up output by 60 to 85%, but if you call a hosted API you pay per token, not per second, so your bill only drops when you control the serving stack or cut the tokens themselves.

Claude Sonnet 5 Pricing: What the Cheaper Agent Model Really Costs
LLM Economics
Jun 30, 20264 min read

Claude Sonnet 5 Pricing: What the Cheaper Agent Model Really Costs

Claude Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens (introductory pricing through August 31, 2026), less than half the price of Opus 4.8, but a new tokenizer and a scheduled rate increase mean your real cost depends on the workload you run.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.