Nebius Pricing
Nebius charges between $0.030 and $4.40 per 1 million output tokens, depending on the model. Qwen Qwen2.5 Coder 7B is the cheapest at $0.030/1M output and $0.010/1M input; GLM-5.2 is the most expensive at $4.40/1M output. The largest context window is 1.0M tokens (MiniMax-M3). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.010/1M
Qwen Qwen2.5 Coder 7B
Cheapest output
$0.030/1M
Qwen Qwen2.5 Coder 7B
Longest context
1.0M
MiniMax-M3
Models priced
58
Avg $1.22/1M output
Pricing by model
All prices in USD per 1 million tokens. All 58 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Qwen Qwen2.5 Coder 7B | $0.010 | $0.030 | 33K |
| Meta Llama Llama Guard 3 8B | $0.020 | $0.060 | 128K |
| Meta Llama Meta Llama 3.1 8B Instruct | $0.020 | $0.060 | 128K |
| Qwen Qwen2 VL 7B Instruct | $0.020 | $0.060 | 131K |
| Mistralai Mistral Nemo Instruct 2407 | $0.040 | $0.120 | 128K |
| Google Gemma 3 27B It | $0.060 | $0.200 | 128K |
| Qwen Qwen2.5 32B Instruct | $0.060 | $0.200 | 128K |
| Nemotron-3-Nano-30B-A3B ToolsCache | $0.060 | $0.240 | 32K |
| Nemotron-3-Nano-Omni ReasoningToolsCache | $0.060 | $0.240 | 66K |
| Qwen Qwen3 14B | $0.080 | $0.240 | 33K |
| Qwen Qwen3 4B | $0.080 | $0.240 | 33K |
| Gemma-3-27b-it VisionToolsCache | $0.100 | $0.300 | 110K |
| Qwen Qwen3 30B A3b | $0.100 | $0.300 | 33K |
| Qwen Qwen3 32B | $0.100 | $0.300 | 33K |
| Qwen3-30B-A3B-Instruct-2507 ToolsCache | $0.100 | $0.300 | 128K |
| Qwen3-32B ToolsCache | $0.100 | $0.300 | 128K |
| Hermes-4-70B ReasoningToolsCache | $0.130 | $0.400 | 128K |
| Llama-3.3-70B-Instruct ToolsCache | $0.130 | $0.400 | 128K |
| Meta Llama Llama 3.3 70B Instruct | $0.130 | $0.400 | 128K |
| Meta Llama Meta Llama 3.1 70B Instruct | $0.130 | $0.400 | 128K |
| Nvidia Llama 3.3 Nemotron Super 49B | $0.100 | $0.400 | 131K |
| Qwen Qwen2 VL 72B Instruct | $0.130 | $0.400 | 131K |
| Qwen Qwen2.5 72B Instruct | $0.130 | $0.400 | 128K |
| Qwen Qwen2.5 VL 72B Instruct | $0.130 | $0.400 | 131K |
| DeepSeek-V3.2 ReasoningToolsCache | $0.300 | $0.450 | 163K |
| Qwen Qwq 32B | $0.150 | $0.450 | 33K |
| gpt-oss-120b-fast ReasoningToolsCache | $0.100 | $0.500 | 8K |
| gpt-oss-120b ReasoningToolsCache | $0.150 | $0.600 | 128K |
| Qwen Qwen3 235B A22b | $0.200 | $0.600 | 262K |
| Qwen3 235B A22B Instruct 2507 Tools | $0.200 | $0.600 | 262K |
| DeepSeek Ai DeepSeek R1 Distill Llama 70B | $0.250 | $0.750 | 128K |
| Qwen2.5-VL-72B-Instruct VisionToolsCache | $0.250 | $0.750 | 128K |
| Nemotron-3-Super-120B-A12B ReasoningTools | $0.300 | $0.900 | 256K |
| INTELLECT-3 ToolsCache | $0.200 | $1.10 | 128K |
| MiniMax-M2.5 ReasoningToolsCache | $0.300 | $1.20 | 197K |
| MiniMax-M2.5-fast ReasoningToolsCache | $0.300 | $1.20 | 8K |
| MiniMax-M3 ReasoningTools | $0.300 | $1.20 | 1.0M |
| Qwen3-Next-80B-A3B-Thinking ReasoningToolsCache | $0.150 | $1.20 | 128K |
| Qwen3-Next-80B-A3B-Thinking-fast ReasoningToolsCache | $0.150 | $1.20 | 8K |
| DeepSeek Ai DeepSeek V3 | $0.500 | $1.50 | 128K |
| DeepSeek Ai DeepSeek V3 0324 | $0.500 | $1.50 | 128K |
| Llama-3.1-Nemotron-Ultra-253B-v1 ToolsCache | $0.600 | $1.80 | 128K |
| Nvidia Llama 3.1 Nemotron Ultra 253B | $0.600 | $1.80 | 128K |
| DeepSeek-V3.2-fast ReasoningToolsCache | $0.400 | $2.00 | 8K |
| Qwen3-235B-A22B-Thinking-2507-fast ReasoningToolsCache | $0.500 | $2.00 | 8K |
| DeepSeek Ai DeepSeek R1 | $0.800 | $2.40 | 128K |
| DeepSeek Ai DeepSeek R1 0528 | $0.800 | $2.40 | 164K |
| Kimi-K2.5 VisionReasoningToolsCache | $0.500 | $2.50 | 256K |
| Kimi-K2.5-fast VisionReasoningToolsCache | $0.500 | $2.50 | 256K |
| Hermes-4-405B ReasoningToolsCache | $1.00 | $3.00 | 128K |
| Meta Llama Meta Llama 3.1 405B Instruct | $1.00 | $3.00 | 128K |
| Nousresearch Hermes 3 Llama 3.1 405B | $1.00 | $3.00 | 128K |
| GLM-5 ReasoningToolsCache | $1.00 | $3.20 | 200K |
| DeepSeek V4 Pro ReasoningToolsCache | $1.75 | $3.50 | 1.0M |
| Qwen3.5-397B-A17B ReasoningToolsCache | $0.600 | $3.60 | 262K |
| Qwen3.5-397B-A17B-fast ReasoningToolsCache | $0.600 | $3.60 | 8K |
| Kimi K2.7 Code ReasoningTools | $0.950 | $4.00 | 262K |
| GLM-5.2 ReasoningTools | $1.40 | $4.40 | 432K |
Frequently asked questions
How much does Nebius cost per 1M tokens?
Nebius pricing ranges from $0.030 to $4.40 per 1 million output tokens across 58 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Nebius model?
Qwen Qwen2.5 Coder 7B is the cheapest Nebius model on both axes — $0.010 per 1M input tokens and $0.030 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Nebius context window?
MiniMax-M3 has the largest context window in the Nebius catalog at 1.0M input tokens, with up to 1.0M output tokens per response.
Does Nebius support prompt caching?
Yes. Nemotron-3-Nano-30B-A3B reads cached input at $0.0060 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Nebius models support reasoning or vision?
The Nebius catalog includes 22 reasoning models, 4 vision models, 31 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Nebius bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator