All articles
The Calcaas Blog

All articles

Page 2 of 7

What Cursor's $7 India Plan Teaches AI SaaS Founders About Geographic Pricing
Pricing Strategy
Jul 28, 20264 min read

What Cursor's $7 India Plan Teaches AI SaaS Founders About Geographic Pricing

Cursor just launched a roughly $7-a-month India-only plan, a third of its $20 Pro price, and the trick that makes it sustainable is routing those users to its own cheaper models instead of frontier third-party ones.

Claude Opus 5 vs Fable 5: What the Half-Price Flagship Means for Your LLM Cost Model
LLM Economics
Jul 28, 20264 min read

Claude Opus 5 vs Fable 5: What the Half-Price Flagship Means for Your LLM Cost Model

Anthropic's Claude Opus 5 lands within a few benchmark points of Fable 5 at roughly half the price, and independent testing shows a 20% lower cost per task, a shift worth rerunning your margin math over.

Why the GPU-Hour Sticker Price Isn't What You'll Pay for AI Inference at Scale
LLM Economics
Jul 27, 20265 min read

Why the GPU-Hour Sticker Price Isn't What You'll Pay for AI Inference at Scale

Headline GPU and LLM token prices describe what one unit costs, not what a usable amount of compute costs once you need it delivered together, on time, and at your scale.

Runway's Pivot From 'Best Model' to 'Best Router' Is a Pricing Story
Pricing Strategy
Jul 25, 20264 min read

Runway's Pivot From 'Best Model' to 'Best Router' Is a Pricing Story

Runway just launched a router that picks the cheapest or best generative-media model per request, weeks after swapping its unlimited plans for token-based pricing, a two-part playbook worth studying if your own product depends on someone else's model.

One Script, 26 Tool Calls, 99.2% Cheaper: What 'Code Mode' Really Saves
LLM Economics
Jul 25, 20264 min read

One Script, 26 Tool Calls, 99.2% Cheaper: What 'Code Mode' Really Saves

A production measurement from Agent Swarm shows that letting an agent run one script instead of 26 sequential tool calls cut a real workflow from an estimated $2.44 floor down to about two cents, and the saving dwarfs any difference between LLM providers.

Claude Opus 5's Price Didn't Change, But Your Cost-Per-Task Probably Just Dropped
Product Updates
Jul 25, 20264 min read

Claude Opus 5's Price Didn't Change, But Your Cost-Per-Task Probably Just Dropped

Opus 5 launches at the same $5/$25 per million token price as its predecessor, but Anthropic's own benchmarks suggest it needs fewer tokens and fewer retries to hit the same result, which is the number that actually matters for your margins.

LLM Economics
Jul 25, 20263 min read

Do AI Model Routers Actually Cut Your LLM Bill by Two-Thirds?

A Show HN launch claims routing across open-weight models delivers Fable-level output at roughly one-third the cost, but the real savings depend on your task mix and whether you keep your prompt cache intact.

Why AI Inference Costs Are Becoming the Real Battleground for SaaS Margins
LLM Economics
Jul 24, 20266 min read

Why AI Inference Costs Are Becoming the Real Battleground for SaaS Margins

AI inference, not training, is now the line item that will make or break your SaaS margins, because unlike training it is a per-query operating cost that scales with every user action.

Gemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins
LLM Economics
Jul 23, 20265 min read

Gemini 3.6 Flash Pricing: What $1.50/$7.50 Per Million Tokens Means for Agent Margins

Google's new Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, undercutting 3.5 Flash on price while using roughly 17% fewer output tokens per task, a combination that compounds into a bigger margin gain than the sticker price alone suggests.

Cutting Your Price 5x Won't Kill Your Margin. Your Architecture Will
Pricing Strategy
Jul 18, 20264 min read

Cutting Your Price 5x Won't Kill Your Margin. Your Architecture Will

Aggressive usage-based pricing does not fail because the price is too low; it fails when founders cut the price without first cutting their cost of goods.

The Hidden Token Tax on AI Agents: Tool Bloat
LLM Economics
Jul 18, 20264 min read

The Hidden Token Tax on AI Agents: Tool Bloat

Loading an agent's full tool catalog into context every turn inflates input tokens on every call; retrieving only the two or three tools that matter per turn cut input tokens by up to 85% in one benchmark.

How to Read the True Cost of Your AI Coding Agent
LLM Economics
Jul 18, 20264 min read

How to Read the True Cost of Your AI Coding Agent

Reconstruct a per-session cost floor from your own agent transcripts, then re-price the same tokens on a cheaper model to see what you could save.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.