LLM Economics

LLM Economics

Token math, model selection, and the true unit cost of every generation.

Page 4 of 5

Anthropic's California Claude Discount: What a 50% Price Cut Really Does to Your LLM Costs
LLM Economics
Jun 30, 20265 min read

Anthropic's California Claude Discount: What a 50% Price Cut Really Does to Your LLM Costs

A 50% discount on Claude does not just halve your bill: it changes your effective cost per token, your gross margin, and the breakeven math on every AI feature you ship.

GPT-5.6 Pricing: What Sol, Terra, and Luna Cost per Token
LLM Economics
Jun 28, 20265 min read

GPT-5.6 Pricing: What Sol, Terra, and Luna Cost per Token

OpenAI's GPT-5.6 family arrives in three priced tiers, Sol at $5/$30, Terra at $2.50/$15, and Luna at $1/$6 per 1M input/output tokens, which means your model pick now moves gross margin more than your prompt does.

Custom AI Chips Will Reshape Token Prices: What Builders Should Do Now
LLM Economics
Jun 27, 20264 min read

Custom AI Chips Will Reshape Token Prices: What Builders Should Do Now

Custom silicon from OpenAI, Google, Apple and SpaceX is built to cut inference cost, but that does not guarantee cheaper API prices for you, so model your margins across price scenarios instead of betting on one rate.

GPT-5.6 Pricing Explained: Sol vs Terra vs Luna Cost Breakdown
LLM Economics
Jun 27, 20264 min read

GPT-5.6 Pricing Explained: Sol vs Terra vs Luna Cost Breakdown

GPT-5.6 ships in three tiers, Sol at $5/$30, Terra at $2.50/$15, and Luna at $1/$6 per million tokens, so the cost decision is now about routing each task to the cheapest tier that clears your quality bar.

How to Cut Your LLM API Costs and Protect Your SaaS Margins
LLM Economics
Jun 27, 20265 min read

How to Cut Your LLM API Costs and Protect Your SaaS Margins

A cheaper model can swing gross margin from roughly 30% to 85%, but only if you model your real token mix first: output tokens, not the headline input price, decide your unit economics.

OpenAI's Internal Token Use Grew Up to 56x: What It Means for Your AI Budget
LLM Economics
Jun 26, 20264 min read

OpenAI's Internal Token Use Grew Up to 56x: What It Means for Your AI Budget

OpenAI's own usage data shows median internal output tokens rising as much as 56x since November 2025, a warning that per-seat AI costs can compound far faster than headline price cuts.

1,000x Cheaper AI Inference: What It Would Actually Do to Your Margins
LLM Economics
Jun 25, 20264 min read

1,000x Cheaper AI Inference: What It Would Actually Do to Your Margins

Even a 1,000x cut in inference power costs would reshape AI unit economics, but only the share of your bill that is energy moves at that rate, not hardware, overhead, or provider markup.

OpenAI's Custom Chip and What It Actually Means for Your API Bill
LLM Economics
Jun 24, 20264 min read

OpenAI's Custom Chip and What It Actually Means for Your API Bill

A custom inference chip lowers what it costs OpenAI to serve a token, but your API price only drops if they pass the savings through, so model your own cost per token instead of betting on hardware headlines.

Gemini 3.5 Flash Gets Computer Use: What It Means for Agent Costs
LLM Economics
Jun 24, 20264 min read

Gemini 3.5 Flash Gets Computer Use: What It Means for Agent Costs

Putting agentic computer use in a budget-tier model can cut cost per step, but total agent cost depends on how many steps a task takes, so cheaper per token does not always mean cheaper per job.

GPU Cloud Providers in Europe 2026: The Real Cost of Data Residency
LLM Economics
Jun 23, 20264 min read

GPU Cloud Providers in Europe 2026: The Real Cost of Data Residency

European GPU clouds offer B200 and H200 capacity with EU data residency and sovereignty, but residency usually carries a price premium that you should model as part of cost per token, not treat as a free checkbox.

Custom AI Chips vs NVIDIA in 2026: What It Means for Your Inference Cost
LLM Economics
Jun 23, 20263 min read

Custom AI Chips vs NVIDIA in 2026: What It Means for Your Inference Cost

Hyperscaler custom chips like Trainium, Google TPU, Maia, and Meta MTIA are built to cut the provider's cost of serving AI, but that only lowers your bill if it shows up as a cheaper per-token price or GPU-hour rate.

Oracle Cloud GPU Pricing in 2026: H100 vs H200 vs B200 Per-Hour Cost
LLM Economics
Jun 23, 20263 min read

Oracle Cloud GPU Pricing in 2026: H100 vs H200 vs B200 Per-Hour Cost

Oracle Cloud prices H100, H200, and B200 GPUs at different per-hour rates, but the cheapest choice depends on your model size and utilization, not on which chip is newest.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.