IBM models
IBM watsonx.ai foundation models with public pay-as-you-go per-million token rates.


Granite 4H Small (watsonx.ai)
$0.0636 / $0.265 per 1M tokens

Llama 4 Maverick (watsonx.ai)
$0.371 / $1.484 per 1M tokens

Llama 3.3 70B (watsonx.ai)
$0.7526 / $0.7526 per 1M tokens

Mistral Small 3.1 24B (watsonx.ai)
$0.106 / $0.318 per 1M tokens

Mistral Large (watsonx.ai)
$0.636 / $1.908 per 1M tokens

GPT OSS 120B (watsonx.ai)
$0.159 / $0.636 per 1M tokens
IBM
IBM watsonx.ai foundation models with public pay-as-you-go per-million token rates.
There are 6 published models here. On a 70/30 mix, combined list rates run from $0.12402 to $1.0176 per 1 million tokens, with a median of $0.5035. The largest listed context window in this group is 1,000,000 tokens. The most recent price check in this group is 2026-08-14.
IBM list rates
USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.
| Model | Input / 1M | Output / 1M | Mixed 70/30 | Context | Verified |
|---|---|---|---|---|---|
| Granite 4H Small (watsonx.ai) | $0.0636 | $0.265 | $0.12402 | 128,000 | |
| Mistral Small 3.1 24B (watsonx.ai) | $0.106 | $0.318 | $0.1696 | 128,000 | |
| GPT OSS 120B (watsonx.ai) | $0.159 | $0.636 | $0.3021 | 128,000 | |
| Llama 4 Maverick (watsonx.ai) | $0.371 | $1.484 | $0.7049 | 1,000,000 | |
| Llama 3.3 70B (watsonx.ai) | $0.7526 | $0.7526 | $0.7526 | 128,000 | |
| Mistral Large (watsonx.ai) | $0.636 | $1.908 | $1.0176 | 128,000 |
Questions about IBM rates
Do these IBM figures include discounts and tax?
No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.
When is the more expensive IBM tier worth the extra?
Mistral Large (watsonx.ai) sits near $1.0176 per 1M tokens on a 70/30 mix, versus $0.12402 on Granite 4H Small (watsonx.ai). Use the high tier when retries or long reasoning actually fail on the cheap one.