API pricing

IBM watsonx Pricing

IBM watsonx charges between $0.100 and $2000.00 per 1 million output tokens, depending on the model. Ibm Granite Guardian 3.2 2B is the cheapest at $0.100/1M output and $0.100/1M input; Bigscience Mt0 Xxl 13B is the most expensive at $2000.00/1M output. The largest context window is 131K tokens (Mistralai Mistral Large). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.060/1M

Ibm Granite 4 H Small

Cheapest output

$0.100/1M

Ibm Granite Guardian 3.2 2B

Longest context

131K

Mistralai Mistral Large

Models priced

28

Avg $144.01/1M output

Pricing by model

All prices in USD per 1 million tokens. All 28 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Ibm Granite Guardian 3.2 2B$0.100$0.1008K
Ibm Granite Vision 3.2 2B$0.100$0.1008K
Meta Llama Llama 3.2 1B Instruct$0.100$0.100128K
Meta Llama Llama 3.2 3B Instruct$0.150$0.150128K
Ibm Granite 3 8B Instruct$0.200$0.2008K
Ibm Granite 3.3 8B Instruct$0.200$0.2008K
Ibm Granite Guardian 3.3 8B$0.200$0.2008K
Ibm Granite 4 H Small$0.060$0.25020K
Mistralai Mistral Small 2503$0.100$0.30032K
Mistralai Mistral Small 3.1 24B Instruct 2503$0.100$0.30032K
Meta Llama Llama 3.2 11B Vision Instruct$0.350$0.350128K
Meta Llama Llama Guard 3 11B Vision$0.350$0.350128K
Mistralai Pixtral 12B 2409$0.350$0.350128K
Ibm Granite Ttm 1024 96 R2$0.380$0.380512
Ibm Granite Ttm 1536 96 R2$0.380$0.380512
Ibm Granite Ttm 512 96 R2$0.380$0.380512
Google Flan T5 Xl 3B$0.600$0.6008K
Ibm Granite 13B Chat$0.600$0.6008K
Ibm Granite 13B Instruct$0.600$0.6008K
OpenAI GPT Oss 120B$0.150$0.6008K
Meta Llama Llama 3.3 70B Instruct$0.710$0.710128K
Meta Llama Llama 4 Maverick 17B$0.350$1.40128K
Sdaia Allam 1 13B Instruct$1.80$1.808K
Meta Llama Llama 3.2 90B Vision Instruct$2.00$2.00128K
Mistralai Mistral Large$3.00$10.00131K
Mistralai Mistral Medium 2505$3.00$10.00128K
Bigscience Mt0 Xxl 13B$500.00$2000.008K
Core42 Jais 13B Chat$500.00$2000.008K

Frequently asked questions

How much does IBM watsonx cost per 1M tokens?

IBM watsonx pricing ranges from $0.100 to $2000.00 per 1 million output tokens across 28 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest IBM watsonx model?

Ibm Granite Guardian 3.2 2B is the cheapest IBM watsonx model on output tokens at $0.100 per 1M, while Ibm Granite 4 H Small is cheapest on input at $0.060 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest IBM watsonx context window?

Mistralai Mistral Large has the largest context window in the IBM watsonx catalog at 131K input tokens, with up to 16K output tokens per response.

How do I calculate my actual IBM watsonx bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator