Cerebras Pricing
Cerebras charges between $0.100 and $2.75 per 1 million output tokens, depending on the model. Llama3.1 8B is the cheapest at $0.100/1M output and $0.100/1M input; Z.AI GLM-4.7 is the most expensive at $2.75/1M output. The largest context window is 131K tokens (Gemma 4 31B IT). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.100/1M
Llama3.1 8B
Cheapest output
$0.100/1M
Llama3.1 8B
Longest context
131K
Gemma 4 31B IT
Models priced
8
Avg $1.30/1M output
Pricing by model
All prices in USD per 1 million tokens. All 8 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Llama3.1 8B | $0.100 | $0.100 | 128K |
| Llama3.1 70B | $0.600 | $0.600 | 128K |
| GPT OSS 120B ReasoningTools | $0.350 | $0.750 | 131K |
| Qwen 3 32B | $0.400 | $0.800 | 128K |
| Llama 3.3 70B | $0.850 | $1.20 | 128K |
| Gemma 4 31B IT VisionReasoningTools | $0.990 | $1.49 | 131K |
| Z.AI GLM-4.7 ReasoningToolsCache | $2.25 | $2.75 | 131K |
| Zai Glm 4.6 | $2.25 | $2.75 | 128K |
Frequently asked questions
How much does Cerebras cost per 1M tokens?
Cerebras pricing ranges from $0.100 to $2.75 per 1 million output tokens across 8 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Cerebras model?
Llama3.1 8B is the cheapest Cerebras model on both axes — $0.100 per 1M input tokens and $0.100 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Cerebras context window?
Gemma 4 31B IT has the largest context window in the Cerebras catalog at 131K input tokens, with up to 41K output tokens per response.
Does Cerebras support prompt caching?
Yes. Z.AI GLM-4.7 reads cached input at $2.25 per 1M tokens, against $2.25 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Cerebras models support reasoning or vision?
The Cerebras catalog includes 3 reasoning models, 1 vision models, 3 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Cerebras bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator