LLM Economics
∑

LLM Economics

Token math, model selection, and the true unit cost of every generation.

Page 3 of 8

AI Agent Hosting Cost: Price Per Completed Task, Not Per GPU-Hour
LLM Economics
Aug 17, 20266 min read

AI Agent Hosting Cost: Price Per Completed Task, Not Per GPU-Hour

On the same GPU at the same hourly rate, a ten-step agent loop costs roughly 14x a single chat call per unit of finished work, and most of that gap is context you re-send rather than work you do.

Why a GPU Cloud Rate Card Is a Ceiling, Not a Price
LLM Economics
Aug 17, 20266 min read

Why a GPU Cloud Rate Card Is a Ceiling, Not a Price

CoreWeave's published H100 rate is $6.16 per GPU-hour, but the best committed discount it advertises lands in the same band as a no-commitment marketplace rate, which means a multi-year contract buys capacity certainty rather than savings.

H100 vs A100 Cost: Why a Faster GPU Can Still Raise Your Bill
LLM Economics
Aug 17, 20265 min read

H100 vs A100 Cost: Why a Faster GPU Can Still Raise Your Bill

An H100 SXM5 costs about 2.75x more per hour than an A100 80GB SXM4, so it only saves money when its throughput advantage clears that same ratio.

Latency Now Has a Price Tag: What Faster Inference Is Actually Worth
LLM Economics
Aug 17, 20265 min read

Latency Now Has a Price Tag: What Faster Inference Is Actually Worth

Anthropic charges exactly double for Fast Mode, which gives you a market price for latency and therefore a ceiling on what any speed optimization is worth to your business.

Open Model Adoption vs Attention: What the Download Data Means for Your Cost Model
LLM Economics
Aug 17, 20265 min read

Open Model Adoption vs Attention: What the Download Data Means for Your Cost Model

Hugging Face's summer 2026 data shows the models the field talks about and the models it actually runs are almost entirely different sets, and your cost model should follow the second one.

Claude Opus 5 Pricing: Why the Same $5/$25 Rate Can Still Raise Your Bill
LLM Economics
Aug 17, 20265 min read

Claude Opus 5 Pricing: Why the Same $5/$25 Rate Can Still Raise Your Bill

Opus 5 kept Opus 4.8's per-token price, but thinking now runs by default, which raises the output share of a typical request and pushes your real blended rate above the sticker math.

Why Harness Optimization Beats Model Switching for Cutting LLM Costs
LLM Economics
Aug 17, 20265 min read

Why Harness Optimization Beats Model Switching for Cutting LLM Costs

Changing the scaffolding around your model cut inference cost by an average of 40% in Writer's own research, which was often a more reliable lever than changing the model itself.

What 14X Faster Inference Does to Your AI Cost Model
LLM Economics
Aug 14, 20265 min read

What 14X Faster Inference Does to Your AI Cost Model

Speed is now a purchasable tier rather than a property of your model, which means latency becomes a line item you choose per endpoint instead of a constraint you inherit.

Agent Memory Token Cost: Why Delivery, Not Storage, Decides Your AI Bill
LLM Economics
Aug 13, 20266 min read

Agent Memory Token Cost: Why Delivery, Not Storage, Decides Your AI Bill

Two agent-memory systems with the same lessons and the same accuracy can differ by nearly 7x in tokens per task, because the cost is set by how much context you send at inference, not by how much you keep.

Distillation Just Got 15x Cheaper. Here Is What That Does to Your AI Cost Model
LLM Economics
Aug 11, 20266 min read

Distillation Just Got 15x Cheaper. Here Is What That Does to Your AI Cost Model

A new memory-efficient training method cut peak VRAM in knowledge distillation from 85.2 GiB to 5.45 GiB at 32K context, which moves the volume at which running your own small model beats paying per token.

DeepSeek's API Price Hike: Why Cheap Tokens Were Never a Strategy
LLM Economics
Aug 10, 20265 min read

DeepSeek's API Price Hike: Why Cheap Tokens Were Never a Strategy

DeepSeek announced on August 6, 2026 that it will raise API prices by a relatively large margin, and the smartest response for AI builders is to model their exposure now, before the number lands.

LLM API Pricing Comparison 2026: GPT, Claude, Gemini and DeepSeek by Blended Cost Per Million Tokens
LLM Economics
Aug 8, 20266 min read

LLM API Pricing Comparison 2026: GPT, Claude, Gemini and DeepSeek by Blended Cost Per Million Tokens

The cheapest LLM API in August 2026 is DeepSeek V4-Flash at $0.14/M input and $0.28/M output, but the number that decides your gross margin is the blended rate for your own input:output mix, not any provider's headline input price.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.