SambaNova Pricing
SambaNova charges from $0.080 per 1 million output tokens, depending on the model. Meta Llama 3.2 1B Instruct is the cheapest at $0.080/1M output and $0.040/1M input; its current flagship Qwen2 Audio 7B Instruct costs $100.00/1M output. The largest context window is 197K tokens (Minimax M2.7). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.040/1M
Meta Llama 3.2 1B Instruct
Cheapest output
$0.080/1M
Meta Llama 3.2 1B Instruct
Longest context
197K
Minimax M2.7
Models priced
19
Avg $7.49/1M output
Pricing by model
All prices in USD per 1 million tokens. All 19 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Meta Llama 3.2 1B Instruct | $0.040 | $0.080 | 16K |
| Meta Llama 3.2 3B Instruct | $0.080 | $0.160 | 4K |
| Meta Llama 3.1 8B Instruct | $0.100 | $0.200 | 16K |
| Meta Llama Guard 3 8B | $0.300 | $0.300 | 16K |
| GPT Oss 120B | $0.220 | $0.590 | 131K |
| Llama 4 Scout 17B 16e Instruct | $0.400 | $0.700 | 8K |
| Qwen3 32B | $0.400 | $0.800 | 8K |
| Qwq 32B | $0.500 | $1.00 | 16K |
| Gemma 4 31B It | $0.380 | $1.15 | 131K |
| Meta Llama 3.3 70B Instruct | $0.600 | $1.20 | 131K |
| DeepSeek R1 Distill Llama 70B | $0.700 | $1.40 | 131K |
| Llama 4 Maverick 17B 128e Instruct | $0.630 | $1.80 | 131K |
| Minimax M2.7 | $0.600 | $2.40 | 197K |
| DeepSeek V3 0324 | $3.00 | $4.50 | 33K |
| DeepSeek V3.1 | $3.00 | $4.50 | 131K |
| DeepSeek V3.2 | $3.00 | $4.50 | 33K |
| DeepSeek R1 | $5.00 | $7.00 | 33K |
| Meta Llama 3.1 405B Instruct | $5.00 | $10.00 | 16K |
| Qwen2 Audio 7B Instruct | $0.500 | $100.00 | 4K |
Frequently asked questions
How much does SambaNova cost per 1M tokens?
SambaNova pricing starts at $0.080 per 1 million output tokens (Meta Llama 3.2 1B Instruct) across 19 models, with its current flagship Qwen2 Audio 7B Instruct at $100.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest SambaNova model?
Meta Llama 3.2 1B Instruct is the cheapest SambaNova model on both axes — $0.040 per 1M input tokens and $0.080 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest SambaNova context window?
Minimax M2.7 has the largest context window in the SambaNova catalog at 197K input tokens, with up to 131K output tokens per response.
How do I calculate my actual SambaNova bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator