IBM
Llama 4 Maverick (watsonx.ai) token cost
Meta Llama 4 Maverick 17B FP8 on watsonx.ai for long-context enterprise chat.

Published list rates
Input
$0.371
per 1M tokens
Output
$1.484
per 1M tokens
Context window
1,000,000
tokens
Pricing source: IBM watsonx.ai pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
Three views of Llama 4 Maverick (watsonx.ai): sending only, receiving only, and a 70/30 mix. Same list rates, bigger rows.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.000000371 | $0.000001 | $0.0000007049 |
| 10 | $0.000004 | $0.000015 | $0.000007 |
| 100 | $0.000037 | $0.000148 | $0.00007 |
| 200 | $0.000074 | $0.000297 | $0.000141 |
| 300 | $0.000111 | $0.000445 | $0.000211 |
| 400 | $0.000148 | $0.000594 | $0.000282 |
| 500 | $0.000186 | $0.000742 | $0.000352 |
| 1,000 | $0.000371 | $0.001484 | $0.000705 |
| 2,000 | $0.000742 | $0.002968 | $0.00141 |
| 5,000 | $0.001855 | $0.00742 | $0.003525 |
| 10,000 | $0.00371 | $0.01484 | $0.007049 |
| 25,000 | $0.009275 | $0.0371 | $0.01762 |
| 50,000 | $0.01855 | $0.0742 | $0.03525 |
| 100,000 | $0.0371 | $0.1484 | $0.07049 |
| 250,000 | $0.09275 | $0.371 | $0.17623 |
| 500,000 | $0.1855 | $0.742 | $0.35245 |
| 1,000,000 | $0.371 | $1.484 | $0.7049 |
| 5,000,000 | $1.855 | $7.42 | $3.5245 |
| 10,000,000 | $3.71 | $14.84 | $7.049 |
| 50,000,000 | $18.55 | $74.20 | $35.245 |
| 100,000,000 | $37.10 | $148.40 | $70.49 |
| 500,000,000 | $185.50 | $742.00 | $352.45 |
| 1,000,000,000 | $371.00 | $1,484.00 | $704.90 |
| 10,000,000,000 | $3,710.00 | $14,840.00 | $7,049.00 |
| 100,000,000,000 | $37,100.00 | $148,400.00 | $70,490.00 |
Tokens for a fixed budget
How far a fixed spend goes on Llama 4 Maverick (watsonx.ai) at these list rates.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 26,954 | 6,739 | 14,186 |
| $0.1 | 269,542 | 67,385 | 141,864 |
| $0.5 | 1.35M | 336,927 | 709,320 |
| $1.00 | 2.7M | 673,854 | 1.42M |
| $5.00 | 13.5M | 3.37M | 7.09M |
| $10.00 | 27M | 6.74M | 14.2M |
| $20.00 | 53.9M | 13.5M | 28.4M |
| $25.00 | 67.4M | 16.8M | 35.5M |
| $30.00 | 80.9M | 20.2M | 42.6M |
| $50.00 | 134.8M | 33.7M | 70.9M |
| $100.00 | 269.5M | 67.4M | 141.9M |
| $250.00 | 673.9M | 168.5M | 354.7M |
| $500.00 | 1.35B | 336.9M | 709.3M |
| $1,000.00 | 2.7B | 673.9M | 1.42B |
| $5,000.00 | 13.5B | 3.37B | 7.09B |
| $10,000.00 | 27B | 6.74B | 14.2B |
At these rates
- A workload with 70% input and 30% output costs $0.7049 per 1 million total tokens and $70.49 per 100 million.
- At the same token count, output costs 4 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 1,000,000-token context window with input alone would cost about $0.371, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.