All articles

Founder Guides

Practical playbooks for SaaS founders modeling cost, trials, and margin.

How to Cut Agent Token Spend by 80% Without Changing Models
Founder Guides
Sep 9, 20266 min read

How to Cut Agent Token Spend by 80% Without Changing Models

A practitioner reports cutting token spend on dynamic agent workflows by roughly 80% through prompt and workflow restructuring alone, which points at an uncomfortable truth: most agent spend is not on the task, it is on context you resend by default.

Log-Scale Pricing Charts Are Quietly Wrecking Your Model Choice
Founder Guides
Sep 5, 20266 min read

Log-Scale Pricing Charts Are Quietly Wrecking Your Model Choice

The intelligence-versus-cost chart everyone quotes uses a logarithmic price axis, which compresses a 50x difference into a couple of centimetres and makes wildly different bills look like neighbours.

I Modeled What $20,000 a Month on Devin Actually Buys a Solo Founder
Founder Guides
Aug 25, 20264 min read

I Modeled What $20,000 a Month on Devin Actually Buys a Solo Founder

A solo founder's public breakdown of spending $20,000 in a month on the Devin coding agent is a useful stress test for how any founder should think about AI-agent ROI, not just whether the number sounds high.

LLM VRAM Sizing: Your Context Limit Is a Pricing Decision
Founder Guides
Aug 17, 20266 min read

LLM VRAM Sizing: Your Context Limit Is a Pricing Decision

Llama 4 Scout needs about 231 GB of VRAM at a 1,024-token context and about 3,671 GB at its advertised 10M-token window, so the context length you configure sets your hardware bill more than the model size does.

How to Cut AI Agent Costs Without Downgrading Your Product
Founder Guides
Aug 14, 20265 min read

How to Cut AI Agent Costs Without Downgrading Your Product

The largest cost lever in an AI agent is no longer which model you call, it is where you draw the line between work that needs judgement and work that only needs data moved.

Idle GPUs Are the Most Expensive Line in Your AI Budget
Founder Guides
Aug 2, 20266 min read

Idle GPUs Are the Most Expensive Line in Your AI Budget

An owned GPU bills you by the calendar hour but only earns by the compute hour, so the utilisation number you assume in your build-versus-buy spreadsheet quietly decides whether the whole decision was right.

How to Manage AI Spend in the Agentic Era: A Founder's Playbook
Founder Guides
Jul 16, 20265 min read

How to Manage AI Spend in the Agentic Era: A Founder's Playbook

OpenAI's new enterprise guidance says to stop staring at token prices and start measuring useful work per dollar; for a founder, that shift is the difference between guessing your margins and knowing them.

The AI Cost Crisis Is Self-Inflicted: How to Control LLM Spend Without Killing Adoption
Founder Guides
Jul 14, 20265 min read

The AI Cost Crisis Is Self-Inflicted: How to Control LLM Spend Without Killing Adoption

Most runaway AI bills come from panic-driven defaults rather than workload needs, and a five-step governance framework can pull spend back without slowing adoption.

LLM Cost Optimization for Founders: Finding the 40-60% of Token Spend You Are Wasting
Founder Guides
Jul 14, 20264 min read

LLM Cost Optimization for Founders: Finding the 40-60% of Token Spend You Are Wasting

Field audits cited by TrueFoundry suggest 40-60% of production LLM token budgets go to redundant calls, oversized models, and ungoverned pipelines, and most of that waste can be located with an afternoon of log analysis.

GLM 5.2 vs Opus: Should You Swap Your Coding Model to Cut Costs?
Founder Guides
Jun 26, 20264 min read

GLM 5.2 vs Opus: Should You Swap Your Coding Model to Cut Costs?

Swapping a premium model like Opus for a cheaper open-weight model like GLM 5.2 can cut your AI bill sharply, but only if it clears the quality bar for the specific work you actually run.

How to Stop Your Team From Burning the AI Budget (Without Banning It)
Founder Guides
Jun 24, 20264 min read

How to Stop Your Team From Burning the AI Budget (Without Banning It)

The durable fix is not rationing tokens after the overspend, it is modeling cost per task up front so every team gets a budget tied to real unit economics.

Self-Hosting vs API: When Local LLMs Actually Cost Less
Founder Guides
Jun 23, 20264 min read

Self-Hosting vs API: When Local LLMs Actually Cost Less

Local open models can run inference at near-zero marginal cost when you reuse hardware you already own, but they are rarely truly free once you count electricity, throughput limits, and engineering time.

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.