All articles
Page 2 of 7
LLM EconomicsOne Script, 26 Tool Calls, 99.2% Cheaper: What 'Code Mode' Really Saves
A production measurement from Agent Swarm shows that letting an agent run one script instead of 26 sequential tool calls cut a real workflow from an estimated $2.44 floor down to about two cents, and the saving dwarfs any difference between LLM providers.
Product UpdatesClaude Opus 5's Price Didn't Change, But Your Cost-Per-Task Probably Just Dropped
Opus 5 launches at the same $5/$25 per million token price as its predecessor, but Anthropic's own benchmarks suggest it needs fewer tokens and fewer retries to hit the same result, which is the number that actually matters for your margins.
Do AI Model Routers Actually Cut Your LLM Bill by Two-Thirds?
A Show HN launch claims routing across open-weight models delivers Fable-level output at roughly one-third the cost, but the real savings depend on your task mix and whether you keep your prompt cache intact.
LLM EconomicsWhy AI Inference Costs Are Becoming the Real Battleground for SaaS Margins
AI inference, not training, is now the line item that will make or break your SaaS margins, because unlike training it is a per-query operating cost that scales with every user action.
LLM EconomicsGemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins
Google's new Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, undercutting 3.5 Flash on price while using roughly 17% fewer output tokens per task, a combination that compounds into a bigger margin gain than the sticker price alone suggests.
Pricing StrategyCutting Your Price 5x Won't Kill Your Margin. Your Architecture Will
Aggressive usage-based pricing does not fail because the price is too low; it fails when founders cut the price without first cutting their cost of goods.
LLM EconomicsThe Hidden Token Tax on AI Agents: Tool Bloat
Loading an agent's full tool catalog into context every turn inflates input tokens on every call; retrieving only the two or three tools that matter per turn cut input tokens by up to 85% in one benchmark.
LLM EconomicsHow to Read the True Cost of Your AI Coding Agent
Reconstruct a per-session cost floor from your own agent transcripts, then re-price the same tokens on a cheaper model to see what you could save.
AI Pricing Needs to Fall 90%? Your Margin Floor Decides Who Survives
Palo Alto Networks CEO Nikesh Arora argues enterprise AI pricing must drop roughly 90% as token costs balloon; whether or not the number is exactly right, the only companies that survive a pricing war are the ones that knew their margin floor before it started.
AI Token Budgets Are Coming: What Meta's Per-Engineer Caps Mean for Founders
Instagram head Adam Mosseri says that within a year or two a strong engineer's AI token burn could match their salary, which means token spend is about to be managed like payroll, and founders should apply the same discipline to per-customer burn.
Founder GuidesHow to Manage AI Spend in the Agentic Era: A Founder's Playbook
OpenAI's new enterprise guidance says to stop staring at token prices and start measuring useful work per dollar; for a founder, that shift is the difference between guessing your margins and knowing them.
LLM EconomicsRenting vs Owning AI: When Do Open Models Beat Frontier APIs on Cost?
Companies typically start on frontier APIs and shift work to open models as usage scales, so rent versus own is a break-even calculation driven by volume and task mix, not an ideological choice.
Pricing math, in your inbox.
One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.