Why a GPU Cloud Rate Card Is a Ceiling, Not a Price
CoreWeave's published H100 rate is $6.16 per GPU-hour, but the best committed discount it advertises lands in the same band as a no-commitment marketplace rate, which means a multi-year contract buys capacity certainty rather than savings.
Aug 17, 2026 · 6 min read
Key takeaways
CoreWeave's published rates are $6.16 per GPU-hour for HGX H100, $6.31 for H200, $8.60 for B200, and $10.50 for GB200 NVL72, all sold as fixed-size instances with an 8-GPU minimum except NVL72's 4-GPU bundles.
The advertised committed discount goes up to 60%, which on the H100 rate lands around $2.46 per GPU-hour. That is roughly what a zero-commitment marketplace already charges on-demand.
CoreWeave's CFO told investors it is largely sold out of 2026 capacity with prices increasing across Ampere, Hopper, and Blackwell.
Its backlog reached $104.2 billion by the end of Q2 2026, up 246% year over year, with a weighted-average new-contract length of roughly five years.
The metric that actually decides your bill is cost per GPU-hour divided by utilization, not the headline rate.
What does CoreWeave actually publish?
Every GPU on the current pricing page is sold as a fixed-size instance rather than a single accelerator:
One oddity is worth flagging: B300 has a published spot rate but no published on-demand rate at all. You can apparently get reclaimable capacity through self-serve checkout, but a guaranteed node requires a sales call. That says something about how tight Blackwell supply still is.
Scale GB200 NVL72 to a full rack of 72 GPUs and you are at roughly $756 an hour, about $544,320 for a continuous 30-day month. That is the real number behind renting rack-scale Blackwell.
Why is the rate card a ceiling rather than a price?
Because the tier that generates most of the revenue is not on it. CoreWeave's business runs on multi-year committed contracts, with a weighted-average new-commitment length of approximately five years and a backlog that hit $104.2 billion by the end of Q2 2026, up 246% year over year on quarterly revenue of $2.6 billion.
CFO Nitin Agrawal put the supply position bluntly: the company remains largely sold out of 2026 capacity, with prices increasing across the board from Ampere to Hopper to Blackwell. CEO Mike Intrator said much the same.
Read plainly, the published on-demand and spot rates describe whatever is left over after contract customers are served, in a market where list prices are rising rather than falling. The rate card exists as a negotiation reference point, not as the number most compute transacts at.
There is also a version-control problem. A legacy Classic pricing page still lists individual H100 PCIe at $4.25/hr and HGX H100 at $4.76/hr, materially below the current page's $6.16 per GPU-hour for nominally the same hardware generation. If you are citing a CoreWeave price you saw somewhere, check which page it came from.
Is a five-year contract actually a discount?
Here is the arithmetic that should reframe the whole reserved-versus-on-demand conversation.
CoreWeave advertises up to 60% off on-demand for committed usage. Apply that to $6.16 per GPU-hour and you land around $2.46. That is essentially identical to CoreWeave's own H100 spot rate with no commitment at all, and it sits in the same band as a marketplace on-demand rate of roughly $2.75 per GPU-hour at single-GPU granularity.
So the best outcome of a successful multi-year negotiation is approximate parity with a zero-commitment alternative.
The observation worth carrying forward: you are not buying a discount, you are buying an option on capacity. In a market that is largely sold out, that option can be genuinely valuable, and it should be priced as one. What it is not is a cost-reduction line in your model, and treating it as one will make your forecast wrong in the direction that hurts.
What actually decides your bill?
Utilization, because of the 8-GPU minimum.
Run the honest comparison. A full 8-GPU B200 node for a 720-hour month costs $68.80 x 720, or $49,536, which works out cheaper than paying $9.36 per GPU-hour for eight separate single-GPU instances at $53,913.60. CoreWeave wins that comparison because it amortises across the whole node.
The catch is the word full. Use six of those eight GPUs and CoreWeave still bills all eight, while per-GPU billing scales down with you. The break-even is not a price, it is a utilization threshold, and almost nobody measures theirs before signing.
So the metric to carry into any GPU procurement conversation is effective cost per GPU-hour, meaning list rate divided by your realistic sustained utilization. A $6.16 rate at 90% utilization beats a $2.75 rate at 35%.
One line to keep: never compare rate cards, compare effective cost per unit of work actually delivered.
You can run that comparison across providers at your own utilization at calcaas.com.
Frequently asked questions
What does CoreWeave charge for GPU compute in 2026?
The published rate card lists HGX H100 at $6.16 per GPU-hour, HGX H200 at $6.31, HGX B200 at $8.60, and GB200 NVL72 at $10.50 per GPU-hour. All are sold as fixed-size instances, generally with an 8-GPU minimum, and spot rates run substantially lower at $2.46 for H100 and $4.26 for B200.
How big is CoreWeave's committed-use discount?
CoreWeave advertises up to 60% off on-demand pricing for committed usage but does not publish which commitment length unlocks which tier. Applied to the $6.16 H100 rate, a full 60% discount lands around $2.46 per GPU-hour, which is close to its own spot rate and to no-commitment marketplace on-demand pricing.
Should I sign a reserved GPU contract to save money?
Treat it as buying an option on capacity rather than a discount. In a market where CoreWeave says it is largely sold out of 2026 capacity with prices rising, guaranteed access has real value, but the negotiated rate is likely to land near what an on-demand alternative already charges.
What determines whether a node-based GPU price is actually cheaper?
Utilization. An 8-GPU B200 node at $68.80 an hour costs $49,536 over a 720-hour month, cheaper than eight single GPUs billed separately, but only if all eight are genuinely busy. Below full utilization you still pay for the whole node, so effective cost per GPU-hour is list rate divided by sustained utilization. Place the JSON-LD above inside a <script type="application/ld+json"> tag in the page head.