IBM
Granite 4H Small (watsonx.ai) token cost
IBM Granite 4H Small pay-as-you-go rates on watsonx.ai for enterprise open models.

Published list rates
Input
$0.0636
per 1M tokens
Output
$0.265
per 1M tokens
Context window
128,000
tokens
Pricing source: IBM watsonx.ai pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
Granite 4H Small (watsonx.ai) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.0000000636 | $0.000000265 | $0.000000124 |
| 10 | $0.000000636 | $0.000003 | $0.000001 |
| 100 | $0.000006 | $0.000027 | $0.000012 |
| 200 | $0.000013 | $0.000053 | $0.000025 |
| 300 | $0.000019 | $0.00008 | $0.000037 |
| 400 | $0.000025 | $0.000106 | $0.00005 |
| 500 | $0.000032 | $0.000133 | $0.000062 |
| 1,000 | $0.000064 | $0.000265 | $0.000124 |
| 2,000 | $0.000127 | $0.00053 | $0.000248 |
| 5,000 | $0.000318 | $0.001325 | $0.00062 |
| 10,000 | $0.000636 | $0.00265 | $0.00124 |
| 25,000 | $0.00159 | $0.006625 | $0.003101 |
| 50,000 | $0.00318 | $0.01325 | $0.006201 |
| 100,000 | $0.00636 | $0.0265 | $0.0124 |
| 250,000 | $0.0159 | $0.06625 | $0.031 |
| 500,000 | $0.0318 | $0.1325 | $0.06201 |
| 1,000,000 | $0.0636 | $0.265 | $0.12402 |
| 5,000,000 | $0.318 | $1.325 | $0.6201 |
| 10,000,000 | $0.636 | $2.65 | $1.2402 |
| 50,000,000 | $3.18 | $13.25 | $6.201 |
| 100,000,000 | $6.36 | $26.50 | $12.402 |
| 500,000,000 | $31.80 | $132.50 | $62.01 |
| 1,000,000,000 | $63.60 | $265.00 | $124.02 |
| 10,000,000,000 | $636.00 | $2,650.00 | $1,240.20 |
| 100,000,000,000 | $6,360.00 | $26,500.00 | $12,402.00 |
Tokens for a fixed budget
If you cap spend, this is roughly how many Granite 4H Small (watsonx.ai) tokens that money buys.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 157,233 | 37,736 | 80,632 |
| $0.1 | 1.57M | 377,358 | 806,322 |
| $0.5 | 7.86M | 1.89M | 4.03M |
| $1.00 | 15.7M | 3.77M | 8.06M |
| $5.00 | 78.6M | 18.9M | 40.3M |
| $10.00 | 157.2M | 37.7M | 80.6M |
| $20.00 | 314.5M | 75.5M | 161.3M |
| $25.00 | 393.1M | 94.3M | 201.6M |
| $30.00 | 471.7M | 113.2M | 241.9M |
| $50.00 | 786.2M | 188.7M | 403.2M |
| $100.00 | 1.57B | 377.4M | 806.3M |
| $250.00 | 3.93B | 943.4M | 2.02B |
| $500.00 | 7.86B | 1.89B | 4.03B |
| $1,000.00 | 15.7B | 3.77B | 8.06B |
| $5,000.00 | 78.6B | 18.9B | 40.3B |
| $10,000.00 | 157.2B | 37.7B | 80.6B |
At these rates
- A workload with 70% input and 30% output costs $0.12402 per 1 million total tokens and $12.402 per 100 million.
- At the same token count, output costs 4.17 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 128,000-token context window with input alone would cost about $0.008141, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.