API pricing

OVHcloud Pricing

OVHcloud charges between $0.100 and $4.25 per 1 million output tokens, depending on the model. Llama 3.1 8B Instruct is the cheapest at $0.100/1M output and $0.100/1M input; Qwen3.5-397B-A17B is the most expensive at $4.25/1M output. The largest context window is 262K tokens (Qwen3-Coder-30B-A3B-Instruct). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.050/1M

gpt-oss-20b

Cheapest output

$0.100/1M

Llama 3.1 8B Instruct

Longest context

262K

Qwen3-Coder-30B-A3B-Instruct

Models priced

19

Avg $0.764/1M output

Pricing by model

All prices in USD per 1 million tokens. All 19 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Llama 3.1 8B Instruct$0.100$0.100131K
Mistral-7B-Instruct-v0.3
Tools
$0.110$0.11066K
Mistral-Nemo-Instruct-2407
Tools
$0.140$0.14066K
gpt-oss-20b
ReasoningTools
$0.050$0.180131K
Qwen3.5-9B
VisionReasoningTools
$0.120$0.180262K
Mamba Codestral 7B V0.1$0.190$0.190256K
Qwen3-32B
ReasoningTools
$0.090$0.25033K
Qwen3-Coder-30B-A3B-Instruct
Tools
$0.070$0.260262K
Llava V1.6 Mistral 7B Hf$0.290$0.29032K
Mistral-Small-3.2-24B-Instruct-2506
VisionTools
$0.100$0.310131K
gpt-oss-120b
ReasoningTools
$0.090$0.470131K
Mixtral 8X7B Instruct V0.1$0.630$0.63032K
DeepSeek R1 Distill Llama 70B$0.670$0.670131K
Meta Llama 3 1 70B Instruct$0.670$0.670131K
Meta-Llama-3_3-70B-Instruct
Tools
$0.740$0.740131K
Qwen2.5 Coder 32B Instruct$0.870$0.87032K
Qwen2.5-VL-72B-Instruct
Vision
$1.01$1.0133K
Qwen3.6-27B
VisionReasoningTools
$0.470$3.19262K
Qwen3.5-397B-A17B
VisionReasoningTools
$0.710$4.25262K

Frequently asked questions

How much does OVHcloud cost per 1M tokens?

OVHcloud pricing ranges from $0.100 to $4.25 per 1 million output tokens across 19 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest OVHcloud model?

Llama 3.1 8B Instruct is the cheapest OVHcloud model on output tokens at $0.100 per 1M, while gpt-oss-20b is cheapest on input at $0.050 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest OVHcloud context window?

Qwen3-Coder-30B-A3B-Instruct has the largest context window in the OVHcloud catalog at 262K input tokens, with up to 262K output tokens per response.

Which OVHcloud models support reasoning or vision?

The OVHcloud catalog includes 6 reasoning models, 5 vision models, 11 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual OVHcloud bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator