LLM Economics
Token math, model selection, and the true unit cost of every generation.
Page 2 of 5
LLM EconomicsWhy AI Inference Costs Are Becoming the Real Battleground for SaaS Margins
AI inference, not training, is now the line item that will make or break your SaaS margins, because unlike training it is a per-query operating cost that scales with every user action.
LLM EconomicsGemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins
Google's new Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, undercutting 3.5 Flash on price while using roughly 17% fewer output tokens per task, a combination that compounds into a bigger margin gain than the sticker price alone suggests.
LLM EconomicsThe Hidden Token Tax on AI Agents: Tool Bloat
Loading an agent's full tool catalog into context every turn inflates input tokens on every call; retrieving only the two or three tools that matter per turn cut input tokens by up to 85% in one benchmark.
LLM EconomicsHow to Read the True Cost of Your AI Coding Agent
Reconstruct a per-session cost floor from your own agent transcripts, then re-price the same tokens on a cheaper model to see what you could save.
AI Token Budgets Are Coming: What Meta's Per-Engineer Caps Mean for Founders
Instagram head Adam Mosseri says that within a year or two a strong engineer's AI token burn could match their salary, which means token spend is about to be managed like payroll, and founders should apply the same discipline to per-customer burn.
LLM EconomicsRenting vs Owning AI: When Do Open Models Beat Frontier APIs on Cost?
Companies typically start on frontier APIs and shift work to open models as usage scales, so rent versus own is a break-even calculation driven by volume and task mix, not an ideological choice.
LLM EconomicsToken Pricing Economics: Will LLM Providers Keep Their Pricing Power?
Benedict Evans argues that today's token prices reflect a temporary supply crunch and that every visible market dynamic points toward frontier models becoming commodity infrastructure, which means founders should plan for falling LLM costs rather than assume today's rate cards.
LLM EconomicsPaying Twice for AI: What Nadella's Warning Means for Your LLM Costs
Satya Nadella argues that companies buying proprietary AI pay twice, once in cash for tokens and again in the proprietary knowledge their usage teaches the model, and that second payment should change how you count AI costs.
LLM EconomicsTokenizer Inflation: Why $/1M Token Prices Are Not Comparable Across LLMs
The same file can become up to 73% more tokens on one frontier model than another, so a $/1M token price is only comparable after you adjust for each model's tokenizer.
How to Think About Token Pricing: 4 Mental Models for AI Founders
A recent Hacker News front-page debate about LLM token pricing keeps circling four mental models, and the one you adopt quietly determines how your product should be priced.
LLM EconomicsThe $3 Trillion AI Question and the Only Version of It You Can Answer
Sequoia's David Cahn now estimates 2026 AI infrastructure spending at $1.5 trillion, implying roughly $3 trillion in revenue to justify it, a question no founder can answer at industry scale but every founder must answer at product scale.
Token Economics for AI Budgets: Why Falling LLM Prices Don't Lower Your Bill
The FinOps Foundation argues that tokens are the atomic unit of AI value, and that consumption growth and workload mix, not falling per-token list prices, now determine what organizations actually spend.
Pricing math, in your inbox.
One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.