All articles
Founder Guides

Log-Scale Pricing Charts Are Quietly Wrecking Your Model Choice

The intelligence-versus-cost chart everyone quotes uses a logarithmic price axis, which compresses a 50x difference into a couple of centimetres and makes wildly different bills look like neighbours.

Sep 5, 2026 · 6 min read
Log-Scale Pricing Charts Are Quietly Wrecking Your Model Choice

Key takeaways

  • The widely shared intelligence-vs-cost plot puts price on a log scale, so enormous absolute gaps look small and trivial gaps look meaningful.
  • Your P&L is linear. A chart that is not linear is answering a different question from the one you are asking.
  • Those charts also tend to list open models at datacenter pricing, which flatters closed models against self-hosting.
  • Gemini 3.8 Flash shipped with better coding and agentic performance at the same introductory pricing as 3.7 Flash. Price-flat upgrades are invisible on a cost axis and matter enormously to your bill.
  • Most products do not need frontier intelligence. Fewer still have measured what frontier intelligence costs them per month.

The chart is not lying, it is answering someone else's question

There is a plot that circulates every time a model launches: intelligence score on one axis, cost on the other, showing the cheapest model that reaches each level of capability. It is genuinely useful, and it is being read by the wrong people.

An analysis flagged in TLDR AI this week made the case bluntly: the cost axis is logarithmic, which means a reader cannot see how enormous the price difference is between the cheap models and the heavy ones, and cannot see how inconsequential the differences are among the cheap ones. Same picture, opposite conclusions, depending on which end you live at.

A log axis is the right choice for a researcher scanning the whole frontier at once. It is the wrong choice for a founder sizing a monthly bill, because nobody has ever received a logarithmic invoice.

Why does the log axis matter so much here?

Because on a log scale, equal distance means equal ratio, not equal dollars.

Two points a centimetre apart near the cheap end might differ by a few cents per million tokens, which is noise for almost any workload. Two points the same centimetre apart at the expensive end differ by tens of dollars per million tokens, which is the difference between a healthy gross margin and a subsidised one.

The eye reads those two gaps as identical. Your accounting does not.

The practical effect is that founders over-optimise at the cheap end, where the savings are rounding errors, and under-react at the expensive end, where the money actually is.

The second distortion: open models priced as if you rented them

The same critique raises a subtler point. These charts typically list open models at their hosted datacenter price, which is always expensive relative to running them yourself on hardware you already have.

That is a defensible methodology for an aggregator that needs one comparable number per model. It is a misleading input for a team weighing self-hosting, because it prices the option at its most expensive form and then plots it against closed models at their normal form. If you are seriously evaluating open weights, the chart is not evidence about you.

The upgrade that never shows up on a cost axis

Gemini 3.8 Flash launched with improved coding, agentic and multi-step reasoning performance at the same introductory pricing as 3.7 Flash. Nothing about that moves on a cost axis, and it is one of the most consequential things that can happen to your bill.

Here is why. Better agentic performance usually means fewer retries and fewer wasted tool calls to finish the same job. Fewer tokens per completed task, at an unchanged rate, is a direct reduction in cost per task. It just does not appear anywhere on a chart whose x-axis is price.

This is the general failure of price-axis thinking: the rate is the variable you can read off a page, and tokens-per-task is the variable that actually determines the invoice.

How do you build a chart that is actually useful to you?

Throw away the leaderboard and plot four numbers instead.

  1. 1Pick one real workload. One task your product performs, with realistic input and output token counts, retries included.
  2. 2Compute the absolute monthly bill for each candidate model at your actual volume. Plain dollars, linear axis, no ratios.
  3. 3Sort by that number. You will usually find two or three models are functionally tied and one is an outlier in each direction.
  4. 4Only then look at quality. Ask whether the quality difference between the tied group and the outlier is worth its specific dollar gap for your specific use case.

A worked example, illustrative only: say two models sit visually adjacent on a log chart, one at $0.50 and one at $5.00 per million output tokens. At 200 million output tokens a month that is $100 against $1,000. On the chart they are neighbours. In your bank account one of them is your entire hosting budget.

The uncomfortable conclusion

Most products do not need frontier intelligence. That claim gets treated as contrarian posturing, and it is mostly just arithmetic that nobody has run. The honest test is not whether the frontier model is better, it obviously is. The test is whether it is better by enough to justify the dollar gap at your volume, and you cannot answer that from a chart with a compressed axis.

Takeaway: before you pick a model from someone else's plot, re-plot it in linear dollars at your own volume, because a chart designed to show the frontier is not designed to show your invoice.

If you would rather not rebuild that in a spreadsheet, you can put your own token volumes in and see the absolute monthly figure in the Calcaas LLM cost calculator.

Frequently asked questions

Why do LLM cost comparison charts use a logarithmic scale?

A log scale lets a single chart display models whose prices differ by several orders of magnitude without the cheapest points collapsing into the axis. It is the right choice for surveying the whole market and the wrong choice for estimating an absolute bill.

What is wrong with reading a log-scale price chart as a founder?

On a log axis, equal visual distance means equal ratio rather than equal dollars. Small gaps at the expensive end can represent tens of dollars per million tokens, while identical-looking gaps at the cheap end represent cents. Budgets are linear, so the visual comparison misleads.

Are open-source models really cheaper than the charts suggest?

Comparison charts usually list open models at hosted datacenter pricing, which is their most expensive form. Teams running open weights on hardware they already operate will see a different cost profile, so the plotted figure is not representative of a self-hosted deployment.

How much does Gemini 3.8 Flash cost?

Google launched Gemini 3.8 Flash with improved coding, agentic and multi-step reasoning performance at the same introductory pricing as Gemini 3.7 Flash. Because the rate is unchanged, any reduction in tokens needed per task translates directly into a lower bill.

How should I compare LLM costs for my own product?

Define one representative task with realistic token counts including retries, calculate the absolute monthly cost for each candidate model at your actual volume, and compare those figures in plain dollars on a linear scale before considering quality differences. Place the JSON-LD above inside a <script type="application/ld+json"> tag in the page head. Source: TLDR AI, 3 September 2026 issue, https://tldr.tech/ai/2026-09-03

ShareXLinkedInFacebook

More from the blog

The Margin Memo

Pricing math, in your inbox.

One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.