MoE Cost per Token: Why Active Parameters Rule
For a mixture-of-experts model, compute per token follows active parameters, while the GPUs you must rent follow total parameters. A model with 40B active out of 400B total can be cheaper per token than a dense 70B at high utilization, and more expensive when traffic is thin.

Last updated: October 2026
Key takeaways
- Spheron argues active parameters, not total parameters, drive GPU-hour cost per token for MoE models.
- Active parameters set compute per token. Total parameters set the memory footprint and the minimum GPU count.
- MoE wins when you can keep the GPUs busy and loses when traffic is low.
- Hosted open-weight APIs hide this math and charge per token instead.
- Model your own utilization before choosing self-hosting over an API.
What does MoE cost per token depend on?
A mixture-of-experts model has many expert sub-networks but routes each token through only a few. The parameters used for one token are the active parameters, and they drive the compute, and therefore the GPU time, per token. Spheron applies this to popular open-weight MoE models and flags two situations where the rule of thumb fails. Credit to their explainer, Cost Per Token MoE: Why Active Parameters Set Your GPU Bill, for the framing.
The catch we want to add: all the experts still have to sit in GPU memory. So total parameters decide how many GPUs you need to rent just to load the model, whether or not you are busy.
How does active vs total parameters change the GPU bill?
Say you compare a dense 70B model with an MoE that has 400B total and 40B active parameters. All numbers below are illustrative, to show the mechanism and not to benchmark any real model.
| Setup (illustrative) | GPU node cost per hour | Throughput (tokens per second) | Cost per 1M tokens | |---|---|---|---| | Dense 70B on a small node | $4 | 1,000 | $1.11 | | MoE 40B active, 400B total on a larger node | $16 | 1,750 | $2.54 |
The dense model: $4 per hour divided by 3.6M tokens per hour is $1.11 per 1M tokens. The MoE: it runs about 1.75x faster because its active parameters are 70 divided by 40 smaller, but it needs a node four times as large to hold all 400B parameters, so $16 divided by 6.3M tokens per hour is $2.54.
So in this scenario the bigger-but-sparser model costs more per token. Active parameters did lower the compute, but the memory bill dominated.
When does MoE become cheaper per token?
Utilization is the swing factor. Batching many requests together lets an MoE spread the cost of its big memory footprint across far more tokens, and throughput per node can rise sharply. If your traffic keeps the node saturated all day, the cost per token can fall below a dense model of similar quality. If traffic is spiky or small, you pay for idle expert weights. Our post on GPU cost per million tokens for self-hosting walks through the utilization math in more detail.
Should you self-host an MoE or use an API?
For most early-stage SaaS products the API wins until volume is steady and high. Hosted open-weight providers charge per token and absorb the idle-capacity risk. Compare per-token rates on Together AI and DeepInfra before you rent a single GPU. We compared open-weight and frontier models in open-weight vs frontier LLM cost comparison.
How should you price a product built on an MoE model?
Base your price on cost per token at your real utilization, not the best-case benchmark. Add a 20 to 30 percent buffer for idle time and retries. If your utilization is likely to vary by season or customer, a credit-based plan protects your margin better than a flat unlimited plan.
Takeaway: active parameters tell you the compute, total parameters tell you the hardware floor, and utilization tells you the bill.
To compare your self-hosted estimate with API rates, open the LLM cost calculator.
Frequently asked questions
What is MoE cost per token based on?
Compute per token follows the active parameters used for each token. The number of GPUs you need follows total parameters because every expert must be loaded in memory. Both matter for your bill.
Is a mixture-of-experts model cheaper to run than a dense model?
It can be, if you keep the GPUs busy, because each token uses fewer parameters. With low traffic, the memory needed for all experts can make it more expensive per token.
How do active parameters affect inference cost?
Fewer active parameters mean less compute and faster tokens per GPU. They do not reduce the memory needed to hold the full model.
Should I use an API instead of self-hosting an MoE model?
Use an API until your traffic is steady and high enough to keep rented GPUs busy. Compare per-token rates first, then model your utilization.
FAQ schema (JSON-LD)
Place this inside a <script type="application/ld+json"> tag in the page head.
More from the blog
LLM EconomicsJevons Paradox in AI: Why Cheaper Tokens Mean Bigger Bills
Jevons paradox in AI means that when tokens get cheaper, people use so many more that total spend rises. If price per task halves from $0.10 to $0.05, volume must double just to hold spend flat, and a tripling lifts spend 50 percent.
LLM EconomicsOpenRouter Pricing: What a Gateway Really Costs
OpenRouter pricing is the underlying model rate plus whatever platform fee the gateway adds, so the real question is net cost. Say a 5 percent fee on $10,000 of monthly tokens adds $500, and routing 15 percent of spend to cheaper models saves $1,500, a net gain of $1,000.
LLM EconomicsJev AI Model Cost: Is It Cheaper per Task?
Jev is a new class of AI model from a ChatGPT co-inventor that promises cheaper and faster software intelligence, but no price list is confirmed here. Judge it by cost per successful task: a model at $0.04 per task with 70 percent success beats a $0.10 incumbent only if failures escalate cheaply.
Pricing math, in your inbox.
One short note a week on AI pricing, token economics, and margin. No spam, unsubscribe anytime.