Nscale Pricing
Nscale charges between $0.025 and $0.600 per 1 million output tokens, depending on the model. DeepSeek Ai DeepSeek R1 Distill Llama 8B is the cheapest at $0.025/1M output and $0.025/1M input; Mistralai Mixtral 8X22B Instruct V0.1 is the most expensive at $0.600/1M output. The largest context window is 0 tokens (DeepSeek Ai DeepSeek R1 Distill Llama 70B). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.010/1M
Qwen Qwen2.5 Coder 3B Instruct
Cheapest output
$0.025/1M
DeepSeek Ai DeepSeek R1 Distill Llama 8B
Longest context
0
DeepSeek Ai DeepSeek R1 Distill Llama 70B
Models priced
14
Avg $0.178/1M output
Pricing by model
All prices in USD per 1 million tokens. All 14 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| DeepSeek Ai DeepSeek R1 Distill Llama 8B | $0.025 | $0.025 | 0 |
| Meta Llama Llama 3.1 8B Instruct | $0.030 | $0.030 | 0 |
| Qwen Qwen2.5 Coder 3B Instruct | $0.010 | $0.030 | 0 |
| Qwen Qwen2.5 Coder 7B Instruct | $0.010 | $0.030 | 0 |
| DeepSeek Ai DeepSeek R1 Distill Qwen 14B | $0.070 | $0.070 | 0 |
| DeepSeek Ai DeepSeek R1 Distill Qwen 1.5b | $0.090 | $0.090 | 0 |
| DeepSeek Ai DeepSeek R1 Distill Qwen 32B | $0.150 | $0.150 | 0 |
| DeepSeek Ai DeepSeek R1 Distill Qwen 7B | $0.200 | $0.200 | 0 |
| Meta Llama Llama 3.3 70B Instruct | $0.200 | $0.200 | 0 |
| Qwen Qwen2.5 Coder 32B Instruct | $0.060 | $0.200 | 0 |
| Qwen Qwq 32B | $0.180 | $0.200 | 0 |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.090 | $0.290 | 0 |
| DeepSeek Ai DeepSeek R1 Distill Llama 70B | $0.375 | $0.375 | 0 |
| Mistralai Mixtral 8X22B Instruct V0.1 | $0.600 | $0.600 | 0 |
Frequently asked questions
How much does Nscale cost per 1M tokens?
Nscale pricing ranges from $0.025 to $0.600 per 1 million output tokens across 14 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Nscale model?
DeepSeek Ai DeepSeek R1 Distill Llama 8B is the cheapest Nscale model on output tokens at $0.025 per 1M, while Qwen Qwen2.5 Coder 3B Instruct is cheapest on input at $0.010 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Nscale context window?
DeepSeek Ai DeepSeek R1 Distill Llama 70B has the largest context window in the Nscale catalog at 0 input tokens, with up to 16K output tokens per response.
How do I calculate my actual Nscale bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator