LLM Economics
Token math, model selection, and the true unit cost of every generation.
Page 2 of 8
LLM EconomicsOpenAI's Jalapeño Chip and What Custom Inference Silicon Means for LLM Pricing
OpenAI's custom Jalapeño inference chip beat the current state-of-the-art Nvidia system on tokens served per user and throughput per kilowatt, a result that points toward cheaper inference and more pressure on provider pricing.
LLM EconomicsA 4-bit Model Just Beat Its Full-Precision Original: What That Means for Your Inference Bill
Multiverse Computing's Quantization-Aware Healing technique made a compressed, 4-bit LLM outperform its own full-precision checkpoint on 7 of 9 benchmarks, while using a fraction of the memory and compute per token.
LLM EconomicsGood Enough and 100x Cheaper Is Beating Best-in-Class: The Economics of AI Simulation
A widely cited breakdown claims simulation-based methods trade a 10% quality hit for 100x lower cost and 10,000x faster iteration, and that math should change how you think about "best" model selection.
LLM EconomicsGPT-5.6 in Kiro: What "Better Price-Performance" Actually Means for Your Token Bill
OpenAI says GPT-5.6 improves price-performance for developers inside Kiro, but the number that decides your margin is cost per finished task, not price per token.
LLM EconomicsOpenAI Is Gaining Ground on Anthropic with Business Users: What That Says About Vendor Lock-In
New data shows businesses swinging between OpenAI and Anthropic as each ships new models, a signal that enterprise AI spend is far less sticky than most vendor contracts assume.
LLM EconomicsAI Compute Just Got a Price Index. Here's What It Means for Your LLM Margins
Silicon Data just raised a $30 million Series A to become the reference price for GPU rental, but a standardized compute market won't automatically make the token prices you pay OpenAI, Anthropic, or Google any more predictable.
LLM EconomicsWhy a 500% Spike in Memory Prices Could Quietly Raise Your AI Costs
DRAM and HBM prices have climbed roughly 500% in 12 months, a hardware shock that sits underneath every GPU and inference bill and could slow or reverse the token-price deflation AI builders have gotten used to.
LLM EconomicsModel Routing: Why Cost, Not Capability, Is Driving Enterprise AI Strategy in 2026
Model routing, sending each AI request to the cheapest model that can handle it, has become a core enterprise cost lever because frontier model prices are rising faster than most teams' AI budgets.
LLM EconomicsWhat OpenRouter's $7B Sale Teaches Builders About Margin at the Routing Layer
Stripe's reported $7 billion purchase of OpenRouter values a token-routing business at roughly 50x revenue, and the real story is how a company with no GPUs of its own posted a 70% gross margin.
LLM EconomicsAnthropic's $65B Revenue Run Rate: What It Means for Your Token Pricing
Anthropic's annualized revenue run rate hit $65 billion in July, a sevenfold jump in seven months, and the growth curve says more about enterprise usage-based pricing than about model quality.
LLM EconomicsStripe's Reported $7B OpenRouter Deal: What It Says About LLM Cost Routing
Stripe has reportedly agreed to buy OpenRouter for more than $7 billion, roughly five times the $1.3 billion valuation the model-routing startup raised at in May 2026, which prices model choice as core financial infrastructure rather than a developer convenience.
LLM EconomicsvLLM vs TensorRT-LLM: The 12% Cost Gap and When It Is Worth Paying For
On the same H100, TensorRT-LLM serves a million output tokens for about $0.66 against vLLM's $0.75, and whether that 12% is worth having depends almost entirely on how often you change models.
Pricing math, in your inbox.
One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.