SambaNova Pricing
SambaNova charges between $0.080 and $100.00 per 1 million output tokens, depending on the model. Meta Llama 3.2 1B Instruct is the cheapest at $0.080/1M output and $0.040/1M input; Qwen2 Audio 7B Instruct is the most expensive at $100.00/1M output. The largest context window is 197K tokens (Minimax M2.7). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.040/1M
Meta Llama 3.2 1B Instruct
Cheapest output
$0.080/1M
Meta Llama 3.2 1B Instruct
Longest context
197K
Minimax M2.7
Models priced
19
Avg $7.49/1M output
Pricing by model
All prices in USD per 1 million tokens. All 19 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Meta Llama 3.2 1B Instruct | $0.040 | $0.080 | 16K |
| Meta Llama 3.2 3B Instruct | $0.080 | $0.160 | 4K |
| Meta Llama 3.1 8B Instruct | $0.100 | $0.200 | 16K |
| Meta Llama Guard 3 8B | $0.300 | $0.300 | 16K |
| GPT Oss 120B | $0.220 | $0.590 | 131K |
| Llama 4 Scout 17B 16e Instruct | $0.400 | $0.700 | 8K |
| Qwen3 32B | $0.400 | $0.800 | 8K |
| Qwq 32B | $0.500 | $1.00 | 16K |
| Gemma 4 31B It | $0.380 | $1.15 | 131K |
| Meta Llama 3.3 70B Instruct | $0.600 | $1.20 | 131K |
| DeepSeek R1 Distill Llama 70B | $0.700 | $1.40 | 131K |
| Llama 4 Maverick 17B 128e Instruct | $0.630 | $1.80 | 131K |
| Minimax M2.7 | $0.600 | $2.40 | 197K |
| DeepSeek V3 0324 | $3.00 | $4.50 | 33K |
| DeepSeek V3.1 | $3.00 | $4.50 | 131K |
| DeepSeek V3.2 | $3.00 | $4.50 | 33K |
| DeepSeek R1 | $5.00 | $7.00 | 33K |
| Meta Llama 3.1 405B Instruct | $5.00 | $10.00 | 16K |
| Qwen2 Audio 7B Instruct | $0.500 | $100.00 | 4K |
Frequently asked questions
How much does SambaNova cost per 1M tokens?
SambaNova pricing ranges from $0.080 to $100.00 per 1 million output tokens across 19 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest SambaNova model?
Meta Llama 3.2 1B Instruct is the cheapest SambaNova model on both axes — $0.040 per 1M input tokens and $0.080 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest SambaNova context window?
Minimax M2.7 has the largest context window in the SambaNova catalog at 197K input tokens, with up to 131K output tokens per response.
How do I calculate my actual SambaNova bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator