IBM watsonx Pricing
IBM watsonx charges between $0.100 and $2000.00 per 1 million output tokens, depending on the model. Ibm Granite Guardian 3.2 2B is the cheapest at $0.100/1M output and $0.100/1M input; Bigscience Mt0 Xxl 13B is the most expensive at $2000.00/1M output. The largest context window is 131K tokens (Mistralai Mistral Large). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.060/1M
Ibm Granite 4 H Small
Cheapest output
$0.100/1M
Ibm Granite Guardian 3.2 2B
Longest context
131K
Mistralai Mistral Large
Models priced
28
Avg $144.01/1M output
Pricing by model
All prices in USD per 1 million tokens. All 28 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Ibm Granite Guardian 3.2 2B | $0.100 | $0.100 | 8K |
| Ibm Granite Vision 3.2 2B | $0.100 | $0.100 | 8K |
| Meta Llama Llama 3.2 1B Instruct | $0.100 | $0.100 | 128K |
| Meta Llama Llama 3.2 3B Instruct | $0.150 | $0.150 | 128K |
| Ibm Granite 3 8B Instruct | $0.200 | $0.200 | 8K |
| Ibm Granite 3.3 8B Instruct | $0.200 | $0.200 | 8K |
| Ibm Granite Guardian 3.3 8B | $0.200 | $0.200 | 8K |
| Ibm Granite 4 H Small | $0.060 | $0.250 | 20K |
| Mistralai Mistral Small 2503 | $0.100 | $0.300 | 32K |
| Mistralai Mistral Small 3.1 24B Instruct 2503 | $0.100 | $0.300 | 32K |
| Meta Llama Llama 3.2 11B Vision Instruct | $0.350 | $0.350 | 128K |
| Meta Llama Llama Guard 3 11B Vision | $0.350 | $0.350 | 128K |
| Mistralai Pixtral 12B 2409 | $0.350 | $0.350 | 128K |
| Ibm Granite Ttm 1024 96 R2 | $0.380 | $0.380 | 512 |
| Ibm Granite Ttm 1536 96 R2 | $0.380 | $0.380 | 512 |
| Ibm Granite Ttm 512 96 R2 | $0.380 | $0.380 | 512 |
| Google Flan T5 Xl 3B | $0.600 | $0.600 | 8K |
| Ibm Granite 13B Chat | $0.600 | $0.600 | 8K |
| Ibm Granite 13B Instruct | $0.600 | $0.600 | 8K |
| OpenAI GPT Oss 120B | $0.150 | $0.600 | 8K |
| Meta Llama Llama 3.3 70B Instruct | $0.710 | $0.710 | 128K |
| Meta Llama Llama 4 Maverick 17B | $0.350 | $1.40 | 128K |
| Sdaia Allam 1 13B Instruct | $1.80 | $1.80 | 8K |
| Meta Llama Llama 3.2 90B Vision Instruct | $2.00 | $2.00 | 128K |
| Mistralai Mistral Large | $3.00 | $10.00 | 131K |
| Mistralai Mistral Medium 2505 | $3.00 | $10.00 | 128K |
| Bigscience Mt0 Xxl 13B | $500.00 | $2000.00 | 8K |
| Core42 Jais 13B Chat | $500.00 | $2000.00 | 8K |
Frequently asked questions
How much does IBM watsonx cost per 1M tokens?
IBM watsonx pricing ranges from $0.100 to $2000.00 per 1 million output tokens across 28 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest IBM watsonx model?
Ibm Granite Guardian 3.2 2B is the cheapest IBM watsonx model on output tokens at $0.100 per 1M, while Ibm Granite 4 H Small is cheapest on input at $0.060 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest IBM watsonx context window?
Mistralai Mistral Large has the largest context window in the IBM watsonx catalog at 131K input tokens, with up to 16K output tokens per response.
How do I calculate my actual IBM watsonx bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator