All articles
The Calcaas Blog

All articles

Page 6 of 10

Gemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins
LLM Economics
Jul 23, 20265 min read

Gemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins

Google's new Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, undercutting 3.5 Flash on price while using roughly 17% fewer output tokens per task, a combination that compounds into a bigger margin gain than the sticker price alone suggests.

Cutting Your Price 5x Won't Kill Your Margin. Your Architecture Will
Pricing Strategy
Jul 18, 20264 min read

Cutting Your Price 5x Won't Kill Your Margin. Your Architecture Will

Aggressive usage-based pricing does not fail because the price is too low; it fails when founders cut the price without first cutting their cost of goods.

The Hidden Token Tax on AI Agents: Tool Bloat
LLM Economics
Jul 18, 20264 min read

The Hidden Token Tax on AI Agents: Tool Bloat

Loading an agent's full tool catalog into context every turn inflates input tokens on every call; retrieving only the two or three tools that matter per turn cut input tokens by up to 85% in one benchmark.

How to Read the True Cost of Your AI Coding Agent
LLM Economics
Jul 18, 20264 min read

How to Read the True Cost of Your AI Coding Agent

Reconstruct a per-session cost floor from your own agent transcripts, then re-price the same tokens on a cheaper model to see what you could save.

Pricing Strategy
Jul 16, 20264 min read

AI Pricing Needs to Fall 90%? Your Margin Floor Decides Who Survives

Palo Alto Networks CEO Nikesh Arora argues enterprise AI pricing must drop roughly 90% as token costs balloon; whether or not the number is exactly right, the only companies that survive a pricing war are the ones that knew their margin floor before it started.

LLM Economics
Jul 16, 20264 min read

AI Token Budgets Are Coming: What Meta's Per-Engineer Caps Mean for Founders

Instagram head Adam Mosseri says that within a year or two a strong engineer's AI token burn could match their salary, which means token spend is about to be managed like payroll, and founders should apply the same discipline to per-customer burn.

How to Manage AI Spend in the Agentic Era: A Founder's Playbook
Founder Guides
Jul 16, 20265 min read

How to Manage AI Spend in the Agentic Era: A Founder's Playbook

OpenAI's new enterprise guidance says to stop staring at token prices and start measuring useful work per dollar; for a founder, that shift is the difference between guessing your margins and knowing them.

Renting vs Owning AI: When Do Open Models Beat Frontier APIs on Cost?
LLM Economics
Jul 14, 20264 min read

Renting vs Owning AI: When Do Open Models Beat Frontier APIs on Cost?

Companies typically start on frontier APIs and shift work to open models as usage scales, so rent versus own is a break-even calculation driven by volume and task mix, not an ideological choice.

Token Pricing Economics: Will LLM Providers Keep Their Pricing Power?
LLM Economics
Jul 14, 20265 min read

Token Pricing Economics: Will LLM Providers Keep Their Pricing Power?

Benedict Evans argues that today's token prices reflect a temporary supply crunch and that every visible market dynamic points toward frontier models becoming commodity infrastructure, which means founders should plan for falling LLM costs rather than assume today's rate cards.

Paying Twice for AI: What Nadella's Warning Means for Your LLM Costs
LLM Economics
Jul 14, 20265 min read

Paying Twice for AI: What Nadella's Warning Means for Your LLM Costs

Satya Nadella argues that companies buying proprietary AI pay twice, once in cash for tokens and again in the proprietary knowledge their usage teaches the model, and that second payment should change how you count AI costs.

The AI Cost Crisis Is Self-Inflicted: How to Control LLM Spend Without Killing Adoption
Founder Guides
Jul 14, 20265 min read

The AI Cost Crisis Is Self-Inflicted: How to Control LLM Spend Without Killing Adoption

Most runaway AI bills come from panic-driven defaults rather than workload needs, and a five-step governance framework can pull spend back without slowing adoption.

Tokenizer Inflation: Why $/1M Token Prices Are Not Comparable Across LLMs
LLM Economics
Jul 14, 20265 min read

Tokenizer Inflation: Why $/1M Token Prices Are Not Comparable Across LLMs

The same file can become up to 73% more tokens on one frontier model than another, so a $/1M token price is only comparable after you adjust for each model's tokenizer.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.