The Calcaas Blog

Cost-first thinking for AI pricing.

Pricing frameworks, LLM economics, product updates, and founder playbooks from the team building the AI cost calculator.

81 articles
Distillation Just Got 15x Cheaper. Here Is What That Does to Your AI Cost Model
LLM EconomicsLatest

Distillation Just Got 15x Cheaper. Here Is What That Does to Your AI Cost Model

A new memory-efficient training method cut peak VRAM in knowledge distillation from 85.2 GiB to 5.45 GiB at 32K context, which moves the volume at which running your own small model beats paying per token.

Aug 11, 2026 · 6 min read
ChatGPT Business Premium Seats: 5x the Price for 5x the Usage, and Why That Is Unusual
Pricing Strategy
Aug 11, 20265 min read

ChatGPT Business Premium Seats: 5x the Price for 5x the Usage, and Why That Is Unusual

OpenAI priced its new Premium seat at $125 per user per month against $25 for Standard, exactly 5x for exactly 5x the usage, which is rarer in SaaS pricing than it sounds.

How to Price AI Products: Founder-Owned Pricing, Surprise Bills, and Inference Margins
Pricing Strategy
Aug 10, 20265 min read

How to Price AI Products: Founder-Owned Pricing, Surprise Bills, and Inference Margins

The short answer from operators at Aiven and Stripe: a founder should own pricing personally, treat it as an iterative computation rather than a one-time decision, and only pass inference through at cost when inference is not the value you sell.

DeepSeek's API Price Hike: Why Cheap Tokens Were Never a Strategy
LLM Economics
Aug 10, 20265 min read

DeepSeek's API Price Hike: Why Cheap Tokens Were Never a Strategy

DeepSeek announced on August 6, 2026 that it will raise API prices by a relatively large margin, and the smartest response for AI builders is to model their exposure now, before the number lands.

LLM API Pricing Comparison 2026: GPT, Claude, Gemini and DeepSeek by Blended Cost Per Million Tokens
LLM Economics
Aug 8, 20266 min read

LLM API Pricing Comparison 2026: GPT, Claude, Gemini and DeepSeek by Blended Cost Per Million Tokens

The cheapest LLM API in August 2026 is DeepSeek V4-Flash at $0.14/M input and $0.28/M output, but the number that decides your gross margin is the blended rate for your own input:output mix, not any provider's headline input price.

Open-Weight vs Frontier LLMs: How to Model the Real Cost Difference
LLM Economics
Aug 6, 20265 min read

Open-Weight vs Frontier LLMs: How to Model the Real Cost Difference

Open-weight models are close enough to frontier quality that the switch is now a pricing decision, but the per-token sticker price hides the serving, safety and compliance work the API price was quietly covering.

Inference Engineering Is a Gross Margin Lever, Not an Infra Detail
LLM Economics
Aug 3, 20265 min read

Inference Engineering Is a Gross Margin Lever, Not an Infra Detail

Inference engineering is the discipline of turning model weights into a fast, affordable production API, and the throughput gains it produces land in someone's gross margin: yours if you serve the model, your provider's if you buy tokens.

LLM Economics
Aug 3, 20265 min read

Open-Weight Models and Your LLM Cost Per Token: What DeepSeek V4 Flash Actually Changes

An open-weight model landing within a few points of the frontier does not automatically cut your AI bill: your real cost per token is set by how many tokens you consume, how fast you can serve them, and whether you own the hardware at all.

Idle GPUs Are the Most Expensive Line in Your AI Budget
Founder Guides
Aug 2, 20266 min read

Idle GPUs Are the Most Expensive Line in Your AI Budget

An owned GPU bills you by the calendar hour but only earns by the compute hour, so the utilisation number you assume in your build-versus-buy spreadsheet quietly decides whether the whole decision was right.

LLM Prices Fell 13x in Four Months: What That Actually Does to Your AI Margins
LLM Economics
Aug 2, 20266 min read

LLM Prices Fell 13x in Four Months: What That Actually Does to Your AI Margins

Frontier-level intelligence now costs roughly one-thirteenth what it did four months ago, but a 13x cut in list price only becomes a 13x cut in your COGS if you re-route workloads and hold token consumption flat.

Flat AI Pricing Is a Subsidy: One User Burned $30,983 of Tokens on a $200 Plan
LLM Economics
Aug 1, 20266 min read

Flat AI Pricing Is a Subsidy: One User Burned $30,983 of Tokens on a $200 Plan

Flat AI subscriptions only work while heavy users stay rare: one developer consuming roughly $30,983 of token value on a $200 per month plan is about a 155x gap between what was paid and what it cost to serve.

OpenAI's GPT-5.6 Price Cut: What Luna at $0.20 and Terra at $2 Do to Your Margins
LLM Economics
Jul 31, 20266 min read

OpenAI's GPT-5.6 Price Cut: What Luna at $0.20 and Terra at $2 Do to Your Margins

Starting July 30, 2026, GPT-5.6 Luna costs $0.20 per 1M input tokens and $1.20 per 1M output tokens (an 80% cut), and Terra costs $2 and $12 (a 20% cut), which means the biggest saving for most teams comes from re-routing workloads rather than from the discount itself.

Cheaper Tokens Do Not Equal Better Margins: What GPT-5.6's Efficiency Push Means for Your AI Unit Economics
LLM Economics
Jul 30, 20265 min read

Cheaper Tokens Do Not Equal Better Margins: What GPT-5.6's Efficiency Push Means for Your AI Unit Economics

Provider efficiency gains only reach your P&L if your own architecture captures them: your prompt-cache hit rate, your agent loop length, and your model tiering decide whether a 20% serving-cost cut becomes 20% of margin or nothing at all.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.