OVHcloud Pricing
OVHcloud charges between $0.100 and $4.25 per 1 million output tokens, depending on the model. Llama 3.1 8B Instruct is the cheapest at $0.100/1M output and $0.100/1M input; Qwen3.5-397B-A17B is the most expensive at $4.25/1M output. The largest context window is 262K tokens (Qwen3-Coder-30B-A3B-Instruct). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.050/1M
gpt-oss-20b
Cheapest output
$0.100/1M
Llama 3.1 8B Instruct
Longest context
262K
Qwen3-Coder-30B-A3B-Instruct
Models priced
19
Avg $0.764/1M output
Pricing by model
All prices in USD per 1 million tokens. All 19 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Llama 3.1 8B Instruct | $0.100 | $0.100 | 131K |
| Mistral-7B-Instruct-v0.3 Tools | $0.110 | $0.110 | 66K |
| Mistral-Nemo-Instruct-2407 Tools | $0.140 | $0.140 | 66K |
| gpt-oss-20b ReasoningTools | $0.050 | $0.180 | 131K |
| Qwen3.5-9B VisionReasoningTools | $0.120 | $0.180 | 262K |
| Mamba Codestral 7B V0.1 | $0.190 | $0.190 | 256K |
| Qwen3-32B ReasoningTools | $0.090 | $0.250 | 33K |
| Qwen3-Coder-30B-A3B-Instruct Tools | $0.070 | $0.260 | 262K |
| Llava V1.6 Mistral 7B Hf | $0.290 | $0.290 | 32K |
| Mistral-Small-3.2-24B-Instruct-2506 VisionTools | $0.100 | $0.310 | 131K |
| gpt-oss-120b ReasoningTools | $0.090 | $0.470 | 131K |
| Mixtral 8X7B Instruct V0.1 | $0.630 | $0.630 | 32K |
| DeepSeek R1 Distill Llama 70B | $0.670 | $0.670 | 131K |
| Meta Llama 3 1 70B Instruct | $0.670 | $0.670 | 131K |
| Meta-Llama-3_3-70B-Instruct Tools | $0.740 | $0.740 | 131K |
| Qwen2.5 Coder 32B Instruct | $0.870 | $0.870 | 32K |
| Qwen2.5-VL-72B-Instruct Vision | $1.01 | $1.01 | 33K |
| Qwen3.6-27B VisionReasoningTools | $0.470 | $3.19 | 262K |
| Qwen3.5-397B-A17B VisionReasoningTools | $0.710 | $4.25 | 262K |
Frequently asked questions
How much does OVHcloud cost per 1M tokens?
OVHcloud pricing ranges from $0.100 to $4.25 per 1 million output tokens across 19 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest OVHcloud model?
Llama 3.1 8B Instruct is the cheapest OVHcloud model on output tokens at $0.100 per 1M, while gpt-oss-20b is cheapest on input at $0.050 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest OVHcloud context window?
Qwen3-Coder-30B-A3B-Instruct has the largest context window in the OVHcloud catalog at 262K input tokens, with up to 262K output tokens per response.
Which OVHcloud models support reasoning or vision?
The OVHcloud catalog includes 6 reasoning models, 5 vision models, 11 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual OVHcloud bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator