All articles

LLM Economics

Token math, model selection, and the true unit cost of every generation.

Open-Weight vs Frontier LLMs: How to Model the Real Cost Difference
LLM Economics
Aug 6, 20265 min read

Open-Weight vs Frontier LLMs: How to Model the Real Cost Difference

Open-weight models are close enough to frontier quality that the switch is now a pricing decision, but the per-token sticker price hides the serving, safety and compliance work the API price was quietly covering.

Inference Engineering Is a Gross Margin Lever, Not an Infra Detail
LLM Economics
Aug 3, 20265 min read

Inference Engineering Is a Gross Margin Lever, Not an Infra Detail

Inference engineering is the discipline of turning model weights into a fast, affordable production API, and the throughput gains it produces land in someone's gross margin: yours if you serve the model, your provider's if you buy tokens.

LLM Economics
Aug 3, 20265 min read

Open-Weight Models and Your LLM Cost Per Token: What DeepSeek V4 Flash Actually Changes

An open-weight model landing within a few points of the frontier does not automatically cut your AI bill: your real cost per token is set by how many tokens you consume, how fast you can serve them, and whether you own the hardware at all.

LLM Prices Fell 13x in Four Months: What That Actually Does to Your AI Margins
LLM Economics
Aug 2, 20266 min read

LLM Prices Fell 13x in Four Months: What That Actually Does to Your AI Margins

Frontier-level intelligence now costs roughly one-thirteenth what it did four months ago, but a 13x cut in list price only becomes a 13x cut in your COGS if you re-route workloads and hold token consumption flat.

Flat AI Pricing Is a Subsidy: One User Burned $30,983 of Tokens on a $200 Plan
LLM Economics
Aug 1, 20266 min read

Flat AI Pricing Is a Subsidy: One User Burned $30,983 of Tokens on a $200 Plan

Flat AI subscriptions only work while heavy users stay rare: one developer consuming roughly $30,983 of token value on a $200 per month plan is about a 155x gap between what was paid and what it cost to serve.

OpenAI's GPT-5.6 Price Cut: What Luna at $0.20 and Terra at $2 Do to Your Margins
LLM Economics
Jul 31, 20266 min read

OpenAI's GPT-5.6 Price Cut: What Luna at $0.20 and Terra at $2 Do to Your Margins

Starting July 30, 2026, GPT-5.6 Luna costs $0.20 per 1M input tokens and $1.20 per 1M output tokens (an 80% cut), and Terra costs $2 and $12 (a 20% cut), which means the biggest saving for most teams comes from re-routing workloads rather than from the discount itself.

Cheaper Tokens Do Not Equal Better Margins: What GPT-5.6's Efficiency Push Means for Your AI Unit Economics
LLM Economics
Jul 30, 20265 min read

Cheaper Tokens Do Not Equal Better Margins: What GPT-5.6's Efficiency Push Means for Your AI Unit Economics

Provider efficiency gains only reach your P&L if your own architecture captures them: your prompt-cache hit rate, your agent loop length, and your model tiering decide whether a 20% serving-cost cut becomes 20% of margin or nothing at all.

Claude Opus 5 vs Fable 5: What the Half-Price Flagship Means for Your LLM Cost Model
LLM Economics
Jul 28, 20264 min read

Claude Opus 5 vs Fable 5: What the Half-Price Flagship Means for Your LLM Cost Model

Anthropic's Claude Opus 5 lands within a few benchmark points of Fable 5 at roughly half the price, and independent testing shows a 20% lower cost per task, a shift worth rerunning your margin math over.

Why the GPU-Hour Sticker Price Isn't What You'll Pay for AI Inference at Scale
LLM Economics
Jul 27, 20265 min read

Why the GPU-Hour Sticker Price Isn't What You'll Pay for AI Inference at Scale

Headline GPU and LLM token prices describe what one unit costs, not what a usable amount of compute costs once you need it delivered together, on time, and at your scale.

One Script, 26 Tool Calls, 99.2% Cheaper: What 'Code Mode' Really Saves
LLM Economics
Jul 25, 20264 min read

One Script, 26 Tool Calls, 99.2% Cheaper: What 'Code Mode' Really Saves

A production measurement from Agent Swarm shows that letting an agent run one script instead of 26 sequential tool calls cut a real workflow from an estimated $2.44 floor down to about two cents, and the saving dwarfs any difference between LLM providers.

LLM Economics
Jul 25, 20263 min read

Do AI Model Routers Actually Cut Your LLM Bill by Two-Thirds?

A Show HN launch claims routing across open-weight models delivers Fable-level output at roughly one-third the cost, but the real savings depend on your task mix and whether you keep your prompt cache intact.

Why AI Inference Costs Are Becoming the Real Battleground for SaaS Margins
LLM Economics
Jul 24, 20266 min read

Why AI Inference Costs Are Becoming the Real Battleground for SaaS Margins

AI inference, not training, is now the line item that will make or break your SaaS margins, because unlike training it is a per-query operating cost that scales with every user action.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.