API pricing

Nebius Pricing

Nebius charges between $0.030 and $4.40 per 1 million output tokens, depending on the model. Qwen Qwen2.5 Coder 7B is the cheapest at $0.030/1M output and $0.010/1M input; GLM-5.2 is the most expensive at $4.40/1M output. The largest context window is 1.0M tokens (MiniMax-M3). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.010/1M

Qwen Qwen2.5 Coder 7B

Cheapest output

$0.030/1M

Qwen Qwen2.5 Coder 7B

Longest context

1.0M

MiniMax-M3

Models priced

58

Avg $1.22/1M output

Pricing by model

All prices in USD per 1 million tokens. All 58 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Qwen Qwen2.5 Coder 7B$0.010$0.03033K
Meta Llama Llama Guard 3 8B$0.020$0.060128K
Meta Llama Meta Llama 3.1 8B Instruct$0.020$0.060128K
Qwen Qwen2 VL 7B Instruct$0.020$0.060131K
Mistralai Mistral Nemo Instruct 2407$0.040$0.120128K
Google Gemma 3 27B It$0.060$0.200128K
Qwen Qwen2.5 32B Instruct$0.060$0.200128K
Nemotron-3-Nano-30B-A3B
ToolsCache
$0.060$0.24032K
Nemotron-3-Nano-Omni
ReasoningToolsCache
$0.060$0.24066K
Qwen Qwen3 14B$0.080$0.24033K
Qwen Qwen3 4B$0.080$0.24033K
Gemma-3-27b-it
VisionToolsCache
$0.100$0.300110K
Qwen Qwen3 30B A3b$0.100$0.30033K
Qwen Qwen3 32B$0.100$0.30033K
Qwen3-30B-A3B-Instruct-2507
ToolsCache
$0.100$0.300128K
Qwen3-32B
ToolsCache
$0.100$0.300128K
Hermes-4-70B
ReasoningToolsCache
$0.130$0.400128K
Llama-3.3-70B-Instruct
ToolsCache
$0.130$0.400128K
Meta Llama Llama 3.3 70B Instruct$0.130$0.400128K
Meta Llama Meta Llama 3.1 70B Instruct$0.130$0.400128K
Nvidia Llama 3.3 Nemotron Super 49B$0.100$0.400131K
Qwen Qwen2 VL 72B Instruct$0.130$0.400131K
Qwen Qwen2.5 72B Instruct$0.130$0.400128K
Qwen Qwen2.5 VL 72B Instruct$0.130$0.400131K
DeepSeek-V3.2
ReasoningToolsCache
$0.300$0.450163K
Qwen Qwq 32B$0.150$0.45033K
gpt-oss-120b-fast
ReasoningToolsCache
$0.100$0.5008K
gpt-oss-120b
ReasoningToolsCache
$0.150$0.600128K
Qwen Qwen3 235B A22b$0.200$0.600262K
Qwen3 235B A22B Instruct 2507
Tools
$0.200$0.600262K
DeepSeek Ai DeepSeek R1 Distill Llama 70B$0.250$0.750128K
Qwen2.5-VL-72B-Instruct
VisionToolsCache
$0.250$0.750128K
Nemotron-3-Super-120B-A12B
ReasoningTools
$0.300$0.900256K
INTELLECT-3
ToolsCache
$0.200$1.10128K
MiniMax-M2.5
ReasoningToolsCache
$0.300$1.20197K
MiniMax-M2.5-fast
ReasoningToolsCache
$0.300$1.208K
MiniMax-M3
ReasoningTools
$0.300$1.201.0M
Qwen3-Next-80B-A3B-Thinking
ReasoningToolsCache
$0.150$1.20128K
Qwen3-Next-80B-A3B-Thinking-fast
ReasoningToolsCache
$0.150$1.208K
DeepSeek Ai DeepSeek V3$0.500$1.50128K
DeepSeek Ai DeepSeek V3 0324$0.500$1.50128K
Llama-3.1-Nemotron-Ultra-253B-v1
ToolsCache
$0.600$1.80128K
Nvidia Llama 3.1 Nemotron Ultra 253B$0.600$1.80128K
DeepSeek-V3.2-fast
ReasoningToolsCache
$0.400$2.008K
Qwen3-235B-A22B-Thinking-2507-fast
ReasoningToolsCache
$0.500$2.008K
DeepSeek Ai DeepSeek R1$0.800$2.40128K
DeepSeek Ai DeepSeek R1 0528$0.800$2.40164K
Kimi-K2.5
VisionReasoningToolsCache
$0.500$2.50256K
Kimi-K2.5-fast
VisionReasoningToolsCache
$0.500$2.50256K
Hermes-4-405B
ReasoningToolsCache
$1.00$3.00128K
Meta Llama Meta Llama 3.1 405B Instruct$1.00$3.00128K
Nousresearch Hermes 3 Llama 3.1 405B$1.00$3.00128K
GLM-5
ReasoningToolsCache
$1.00$3.20200K
DeepSeek V4 Pro
ReasoningToolsCache
$1.75$3.501.0M
Qwen3.5-397B-A17B
ReasoningToolsCache
$0.600$3.60262K
Qwen3.5-397B-A17B-fast
ReasoningToolsCache
$0.600$3.608K
Kimi K2.7 Code
ReasoningTools
$0.950$4.00262K
GLM-5.2
ReasoningTools
$1.40$4.40432K

Frequently asked questions

How much does Nebius cost per 1M tokens?

Nebius pricing ranges from $0.030 to $4.40 per 1 million output tokens across 58 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Nebius model?

Qwen Qwen2.5 Coder 7B is the cheapest Nebius model on both axes — $0.010 per 1M input tokens and $0.030 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Nebius context window?

MiniMax-M3 has the largest context window in the Nebius catalog at 1.0M input tokens, with up to 1.0M output tokens per response.

Does Nebius support prompt caching?

Yes. Nemotron-3-Nano-30B-A3B reads cached input at $0.0060 per 1M tokens, against $0.060 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Nebius models support reasoning or vision?

The Nebius catalog includes 22 reasoning models, 4 vision models, 31 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Nebius bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator