IBM models

IBM watsonx.ai foundation models with public pay-as-you-go per-million token rates.

https://www.ibm.com/products/watsonx-ai/pricing

IBM

IBM watsonx.ai foundation models with public pay-as-you-go per-million token rates.

There are 6 published models here. On a 70/30 mix, combined list rates run from $0.12402 to $1.0176 per 1 million tokens, with a median of $0.5035. The largest listed context window in this group is 1,000,000 tokens. The most recent price check in this group is 2026-08-14.

Official IBM pricing page

IBM list rates

USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.

ModelInput / 1MOutput / 1MMixed 70/30ContextVerified
Granite 4H Small (watsonx.ai) $0.0636 $0.265 $0.12402 128,000
Mistral Small 3.1 24B (watsonx.ai) $0.106 $0.318 $0.1696 128,000
GPT OSS 120B (watsonx.ai) $0.159 $0.636 $0.3021 128,000
Llama 4 Maverick (watsonx.ai) $0.371 $1.484 $0.7049 1,000,000
Llama 3.3 70B (watsonx.ai) $0.7526 $0.7526 $0.7526 128,000
Mistral Large (watsonx.ai) $0.636 $1.908 $1.0176 128,000

Questions about IBM rates

Do these IBM figures include discounts and tax?

No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.

When is the more expensive IBM tier worth the extra?

Mistral Large (watsonx.ai) sits near $1.0176 per 1M tokens on a 70/30 mix, versus $0.12402 on Granite 4H Small (watsonx.ai). Use the high tier when retries or long reasoning actually fail on the cheap one.