Credit Pricing and Margin Guardrails: What Leonardo.ai's Token Model Teaches AI Founders
A credit layer works when it absorbs variable compute cost, and it only stays safe when your unlimited tiers cover models you actually own.
Aug 17, 2026 · 5 min read
Key takeaways
Leonardo.ai's subscription tiers bundle from 8,500 fast tokens a month on Essential up to 60,000 on Ultimate, with unused tokens rolling into a bank capped at 3x monthly tier capacity.
Tokens map to computational intensity, not to assets, so a 4K multi-step render costs more tokens than a small thumbnail.
Unlimited relaxed generation on the $30 Premium and $60 Ultimate tiers applies only to Leonardo's own models. Third-party models always require metered tokens.
Team tiers convert individual token banks into a shared pool with up to 540,000 bank capacity, turning credits into a budgeting tool.
The transferable rule: never underwrite a variable third-party cost inside a flat-price tier.
Why abstract cost into a credit at all?
In text-based AI, tokens are cheap enough that per-request cost variance rarely breaks a plan. In generative media it is the opposite: one click can cost orders of magnitude more than another, and if the output misses the creative brief, the vendor still pays 100% of the render cost.
A credit is a shock absorber between that variance and the price the user sees. It lets you re-weight what a job costs without republishing a price list, and it keeps the customer's mental model simple.
The design mistake to avoid is flat per-asset pricing, for example a fixed price per image. It looks clean and it is margin-negative on exactly the jobs your heaviest users run most.
How does dynamic metering protect margin?
Leonardo meters computational intensity rather than output count. Heavier work, longer sequences, higher resolution, extra passes, consumes more tokens. Light exploratory prompts stay cheap and accessible.
That one decision keeps the billing meter moving in the same direction as the underlying GPU or API bill. If your cost driver is compute and your price driver is asset count, those two curves diverge, and they diverge fastest with your best customers.
What does the rollover bank actually buy you?
Pure consumption models create use-it-or-lose-it anxiety, which shows up as end-of-month churn. Leonardo's answer is an automated rollover bank, capped at 3x the monthly tier capacity.
Note the cap. Uncapped rollover is an open-ended liability sitting on your balance sheet; a 3x ceiling gives users a real safety net while bounding your exposure. Top-ups are handled the same way: they do not expire, but they require an active subscription, which protects the recurring floor.
That is the founder-level lesson. Generosity in a usage model is fine as long as it is bounded generosity.
Why is unlimited only unlimited on first-party models?
This is the sharpest part of the design. On the $30 Premium and $60 Ultimate tiers, users who exhaust their fast tokens can keep generating in a lower-priority relaxed queue, but only on Leonardo's own models such as Phoenix, Lucid Origin, and Motion 2.0. Third-party models like Flux.1, Ideogram, Kling, and Sora 2 always require metered tokens.
The logic is pure unit economics. First-party jobs run on infrastructure Leonardo controls, so they can be queued into off-peak capacity at near-zero marginal cost. Third-party calls carry a hard external bill on every single request. Putting those behind a metered firewall means no customer can ever run up an unbounded partner invoice on a fixed monthly fee.
Here is the observation worth stealing: your unlimited tier is a bet on which costs you can throttle. Any cost you cannot throttle does not belong in it. That rule generalises well beyond generative media, to API resale, data enrichment, and anything you buy per call and sell per seat.
How does this scale into teams?
On team tiers, individual token banks become a shared pool with up to 540,000 bank capacity, bought through a per-seat monthly fee. Credits stop being a personal counter and become a resource-planning instrument: a studio can allocate budget across projects while the vendor keeps a predictable per-seat floor plus platform premiums on collaboration features.
The takeaway: design your credit layer so it flexes with your infrastructure bill, cap every form of generosity, and keep costs you do not control strictly metered.
If you are sizing a credit or hybrid tier of your own, you can model the token, credit, and hybrid variants side by side at calcaas.com.
Frequently asked questions
Why use credits or tokens instead of flat per-asset pricing?
Because the cost of one generation can be orders of magnitude higher than another. Flat per-asset pricing charges the same for a light thumbnail and a heavy multi-step 4K render, so margin collapses on exactly the jobs your heaviest users run most.
Should unused credits roll over?
Rollover reduces the use-it-or-lose-it anxiety that drives end-of-month churn, but it should be capped. Leonardo caps its rollover bank at 3x the monthly tier capacity, which gives users a real safety net while bounding the vendor's outstanding liability.
Can an AI product safely offer an unlimited tier?
Only on capacity it controls. Leonardo's unlimited relaxed generation covers its own first-party models, which it can queue into off-peak infrastructure, while third-party models that carry a hard per-call external bill always require metered tokens.
How do credit models work for teams?
Team tiers replace individual banks with a shared pool, in Leonardo's case up to 540,000 bank capacity, bought through a per-seat monthly fee. That turns credits into a budgeting tool for allocating spend across projects while preserving predictable recurring revenue for the vendor. Place the JSON-LD above inside a <script type="application/ld+json"> tag in the page head.