Claude Opus 5's Price Didn't Change, But Your Cost-Per-Task Probably Just Dropped
Opus 5 launches at the same $5/$25 per million token price as its predecessor, but Anthropic's own benchmarks suggest it needs fewer tokens and fewer retries to hit the same result, which is the number that actually matters for your margins.
Jul 25, 2026 · 4 min read
Key takeaways
Claude Opus 5 launched July 24, 2026, priced at $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8.
Anthropic says Opus 5 comes close to Fable 5's frontier intelligence at about half the price, and on CursorBench it lands within 0.5% of Fable 5's peak score at half the cost per task.
The sticker price staying flat while quality improves means your effective cost per completed task can fall even without a pricing change, if the model needs fewer retries or shorter reasoning chains to succeed.
Opus 5 is also available in a Fast mode at 2x the base price for roughly 2.5x the speed, a separate lever for teams trading cost against latency.
If you're modeling margins on Claude usage, track cost-per-successful-task, not just $/1M tokens, since token price is only half the equation.
What actually changed with Claude Opus 5's pricing?
Nothing, on paper: Opus 5 is priced identically to Opus 4.8, at $5 per million input tokens and $25 per million output tokens. What changed is what you get for that price. Anthropic reports the model performing close to Fable 5, its more expensive frontier model, on coding and knowledge-work benchmarks, in some cases within half a percentage point of Fable 5's best score at roughly half the per-task cost.
Why does a flat token price still change your margins?
Because $/1M tokens is only one input into your real unit economics. If Opus 5 needs fewer turns, fewer retries, or a shorter reasoning chain to complete the same task as Opus 4.8, your total tokens consumed per completed job drops even though the price per token is unchanged. For example, if a task that used to take three attempts at Opus 4.8 now succeeds in one attempt at Opus 5, your effective cost per successful output could fall by more than half, without Anthropic ever touching the price list.
What about Fast mode?
Opus 5 also ships with a Fast mode running at roughly 2.5x the default speed, priced at twice Opus 5's base rate. That's a separate dial: pay more per token to cut latency, which matters if your product is latency-sensitive (a live chat feature, say) but may not be worth it for an async batch job where you can absorb the extra seconds.
How should you model this in your own pricing?
Don't take $/1M tokens as the whole story. Track cost per completed task (or cost per resolved ticket, cost per generated report, whatever your unit of value is) across a model upgrade, because that's the number that flows into your margin, not the sticker price. This is exactly the kind of before-and-after scenario worth running in Calcaas when a new model drops: same price list, different real cost.
A flat token price can still mean a real margin improvement if the new model needs less work to finish the job, so model cost-per-task, not just cost-per-token, and you can run that comparison in Calcaas before you migrate.
Frequently asked questions
Did Claude Opus 5 get more expensive than Opus 4.8?
No. Opus 5 is priced at $5 per million input tokens and $25 per million output tokens, the same rate as Opus 4.8.
Is Opus 5 as good as Fable 5?
Anthropic describes Opus 5 as close to Fable 5's frontier intelligence at roughly half the price, and on one coding benchmark (CursorBench) it scored within 0.5% of Fable 5's peak result at half the cost per task, though it still trails on some evaluations like cybersecurity tasks.
What is Fast mode and does it cost more?
Fast mode runs Opus 5 at about 2.5x the default speed, at twice the base per-token price, useful for latency-sensitive products.
Why should I track cost-per-task instead of just $/1M tokens?
Because token price alone doesn't capture how many tokens or retries a task actually needs. A cheaper-per-token model that requires more attempts can cost more per completed task than a pricier model that succeeds in fewer tries. (Place the JSON-LD block above inside a <script type="application/ld+json"> tag in the page head.)