Anyscale Pricing
Anyscale charges from $0.150 per 1 million output tokens, depending on the model. Google Gemma 7B It is the cheapest at $0.150/1M output and $0.150/1M input; its current flagship Codellama Codellama 34B Instruct Hf costs $1.00/1M output. The largest context window is 66K tokens (Mistralai Mixtral 8X22B Instruct V0.1). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.150/1M
Google Gemma 7B It
Cheapest output
$0.150/1M
Google Gemma 7B It
Longest context
66K
Mistralai Mixtral 8X22B Instruct V0.1
Models priced
12
Avg $0.504/1M output
Pricing by model
All prices in USD per 1 million tokens. All 12 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Google Gemma 7B It | $0.150 | $0.150 | 8K |
| Huggingfaceh4 Zephyr 7B Beta | $0.150 | $0.150 | 16K |
| Meta Llama Llama 2 7B Chat Hf | $0.150 | $0.150 | 4K |
| Meta Llama Meta Llama 3 8B Instruct | $0.150 | $0.150 | 8K |
| Mistralai Mistral 7B Instruct V0.1 | $0.150 | $0.150 | 16K |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.150 | $0.150 | 16K |
| Meta Llama Llama 2 13B Chat Hf | $0.250 | $0.250 | 4K |
| Mistralai Mixtral 8X22B Instruct V0.1 | $0.900 | $0.900 | 66K |
| Codellama Codellama 34B Instruct Hf | $1.00 | $1.00 | 4K |
| Codellama Codellama 70B Instruct Hf | $1.00 | $1.00 | 4K |
| Meta Llama Llama 2 70B Chat Hf | $1.00 | $1.00 | 4K |
| Meta Llama Meta Llama 3 70B Instruct | $1.00 | $1.00 | 8K |
Frequently asked questions
How much does Anyscale cost per 1M tokens?
Anyscale pricing starts at $0.150 per 1 million output tokens (Google Gemma 7B It) across 12 models, with its current flagship Codellama Codellama 34B Instruct Hf at $1.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Anyscale model?
Google Gemma 7B It is the cheapest Anyscale model on both axes — $0.150 per 1M input tokens and $0.150 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Anyscale context window?
Mistralai Mixtral 8X22B Instruct V0.1 has the largest context window in the Anyscale catalog at 66K input tokens, with up to 66K output tokens per response.
How do I calculate my actual Anyscale bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator