LLM Economics

LLM Economics

Token math, model selection, and the true unit cost of every generation.

Page 2 of 8

OpenAI's Jalapeño Chip and What Custom Inference Silicon Means for LLM Pricing
LLM Economics
Aug 26, 20264 min read

OpenAI's Jalapeño Chip and What Custom Inference Silicon Means for LLM Pricing

OpenAI's custom Jalapeño inference chip beat the current state-of-the-art Nvidia system on tokens served per user and throughput per kilowatt, a result that points toward cheaper inference and more pressure on provider pricing.

A 4-bit Model Just Beat Its Full-Precision Original: What That Means for Your Inference Bill
LLM Economics
Aug 26, 20264 min read

A 4-bit Model Just Beat Its Full-Precision Original: What That Means for Your Inference Bill

Multiverse Computing's Quantization-Aware Healing technique made a compressed, 4-bit LLM outperform its own full-precision checkpoint on 7 of 9 benchmarks, while using a fraction of the memory and compute per token.

Good Enough and 100x Cheaper Is Beating Best-in-Class: The Economics of AI Simulation
LLM Economics
Aug 25, 20264 min read

Good Enough and 100x Cheaper Is Beating Best-in-Class: The Economics of AI Simulation

A widely cited breakdown claims simulation-based methods trade a 10% quality hit for 100x lower cost and 10,000x faster iteration, and that math should change how you think about "best" model selection.

GPT-5.6 in Kiro: What "Better Price-Performance" Actually Means for Your Token Bill
LLM Economics
Aug 25, 20264 min read

GPT-5.6 in Kiro: What "Better Price-Performance" Actually Means for Your Token Bill

OpenAI says GPT-5.6 improves price-performance for developers inside Kiro, but the number that decides your margin is cost per finished task, not price per token.

OpenAI Is Gaining Ground on Anthropic with Business Users: What That Says About Vendor Lock-In
LLM Economics
Aug 25, 20264 min read

OpenAI Is Gaining Ground on Anthropic with Business Users: What That Says About Vendor Lock-In

New data shows businesses swinging between OpenAI and Anthropic as each ships new models, a signal that enterprise AI spend is far less sticky than most vendor contracts assume.

AI Compute Just Got a Price Index. Here's What It Means for Your LLM Margins
LLM Economics
Aug 21, 20264 min read

AI Compute Just Got a Price Index. Here's What It Means for Your LLM Margins

Silicon Data just raised a $30 million Series A to become the reference price for GPU rental, but a standardized compute market won't automatically make the token prices you pay OpenAI, Anthropic, or Google any more predictable.

Why a 500% Spike in Memory Prices Could Quietly Raise Your AI Costs
LLM Economics
Aug 20, 20264 min read

Why a 500% Spike in Memory Prices Could Quietly Raise Your AI Costs

DRAM and HBM prices have climbed roughly 500% in 12 months, a hardware shock that sits underneath every GPU and inference bill and could slow or reverse the token-price deflation AI builders have gotten used to.

Model Routing: Why Cost, Not Capability, Is Driving Enterprise AI Strategy in 2026
LLM Economics
Aug 20, 20265 min read

Model Routing: Why Cost, Not Capability, Is Driving Enterprise AI Strategy in 2026

Model routing, sending each AI request to the cheapest model that can handle it, has become a core enterprise cost lever because frontier model prices are rising faster than most teams' AI budgets.

What OpenRouter's $7B Sale Teaches Builders About Margin at the Routing Layer
LLM Economics
Aug 19, 20264 min read

What OpenRouter's $7B Sale Teaches Builders About Margin at the Routing Layer

Stripe's reported $7 billion purchase of OpenRouter values a token-routing business at roughly 50x revenue, and the real story is how a company with no GPUs of its own posted a 70% gross margin.

Anthropic's $65B Revenue Run Rate: What It Means for Your Token Pricing
LLM Economics
Aug 19, 20264 min read

Anthropic's $65B Revenue Run Rate: What It Means for Your Token Pricing

Anthropic's annualized revenue run rate hit $65 billion in July, a sevenfold jump in seven months, and the growth curve says more about enterprise usage-based pricing than about model quality.

Stripe's Reported $7B OpenRouter Deal: What It Says About LLM Cost Routing
LLM Economics
Aug 18, 20266 min read

Stripe's Reported $7B OpenRouter Deal: What It Says About LLM Cost Routing

Stripe has reportedly agreed to buy OpenRouter for more than $7 billion, roughly five times the $1.3 billion valuation the model-routing startup raised at in May 2026, which prices model choice as core financial infrastructure rather than a developer convenience.

vLLM vs TensorRT-LLM: The 12% Cost Gap and When It Is Worth Paying For
LLM Economics
Aug 17, 20265 min read

vLLM vs TensorRT-LLM: The 12% Cost Gap and When It Is Worth Paying For

On the same H100, TensorRT-LLM serves a million output tokens for about $0.66 against vLLM's $0.75, and whether that 12% is worth having depends almost entirely on how often you change models.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.