All articles
Page 2 of 11
Founder GuidesLog-Scale Pricing Charts Are Quietly Wrecking Your Model Choice
The intelligence-versus-cost chart everyone quotes uses a logarithmic price axis, which compresses a 50x difference into a couple of centimetres and makes wildly different bills look like neighbours.
Pricing StrategyA Frontier Model Launch Is a Repricing Event: A Six-Step Checklist for Founders
When a provider ships a new flagship at a new rate, your cost of goods sold changes before your pricing page does, so treat launch day as a margin review rather than a feature announcement.
LLM EconomicsCost Per Token vs Cost Per Task: Why LLM Pricing Comparisons Keep Misleading Founders
A model priced 2.5x higher per token can still produce a smaller bill, because what you actually pay for is tokens consumed per task, not the rate on the pricing page.
Pricing StrategyMeta Just Put a Price on Your Prompts, and It Is Not 95%
Meta's Muse Spark contributor tier cuts input tokens by 92% and output tokens by 95.3%, so the discount you actually capture depends entirely on your input to output ratio.
LLM EconomicsCompute Capacity Belongs in Your AI Cost Model, Not Just Your News Feed
When a provider adds capacity and raises your rate limits, your effective cost per user can fall even though the list price per token has not moved at all.
LLM EconomicsA Cheaper Model Does Not Lower Your AI Bill: What Fable 5.1 Actually Changes
Anthropic's Fable 5.1 is built to reduce token cost, but your monthly AI bill is set by cost per completed task, not by the number on the model card.
Pricing StrategyOpenAI Is Selling Ads on ChatGPT's Free Tier: What It Means for Your AI Pricing Model
OpenAI has started showing ads to ChatGPT's Free and Go users in India, turning ad revenue into a third monetization lever alongside subscriptions and API usage, a move most AI builders can't copy but should still learn from.
LLM EconomicsOpenAI's Jalapeño Chip and What Custom Inference Silicon Means for LLM Pricing
OpenAI's custom Jalapeño inference chip beat the current state-of-the-art Nvidia system on tokens served per user and throughput per kilowatt, a result that points toward cheaper inference and more pressure on provider pricing.
LLM EconomicsA 4-bit Model Just Beat Its Full-Precision Original: What That Means for Your Inference Bill
Multiverse Computing's Quantization-Aware Healing technique made a compressed, 4-bit LLM outperform its own full-precision checkpoint on 7 of 9 benchmarks, while using a fraction of the memory and compute per token.
Founder GuidesI Modeled What $20,000 a Month on Devin Actually Buys a Solo Founder
A solo founder's public breakdown of spending $20,000 in a month on the Devin coding agent is a useful stress test for how any founder should think about AI-agent ROI, not just whether the number sounds high.
LLM EconomicsGood Enough and 100x Cheaper Is Beating Best-in-Class: The Economics of AI Simulation
A widely cited breakdown claims simulation-based methods trade a 10% quality hit for 100x lower cost and 10,000x faster iteration, and that math should change how you think about "best" model selection.
LLM EconomicsGPT-5.6 in Kiro: What "Better Price-Performance" Actually Means for Your Token Bill
OpenAI says GPT-5.6 improves price-performance for developers inside Kiro, but the number that decides your margin is cost per finished task, not price per token.
Pricing math, in your inbox.
One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.