IBM
Llama 3.3 70B (watsonx.ai) token cost
Llama 3.3 70B Instruct on watsonx.ai at the published pay-as-you-go token rate.

Published list rates
Input
$0.7526
per 1M tokens
Output
$0.7526
per 1M tokens
Context window
128,000
tokens
Pricing source: IBM watsonx.ai pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
What Llama 3.3 70B (watsonx.ai) costs if you only send tokens, only get a reply, or do both (70% in / 30% out). Rows start small and go up to huge traffic.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.0000007526 | $0.0000007526 | $0.0000007526 |
| 10 | $0.000008 | $0.000008 | $0.000008 |
| 100 | $0.000075 | $0.000075 | $0.000075 |
| 200 | $0.000151 | $0.000151 | $0.000151 |
| 300 | $0.000226 | $0.000226 | $0.000226 |
| 400 | $0.000301 | $0.000301 | $0.000301 |
| 500 | $0.000376 | $0.000376 | $0.000376 |
| 1,000 | $0.000753 | $0.000753 | $0.000753 |
| 2,000 | $0.001505 | $0.001505 | $0.001505 |
| 5,000 | $0.003763 | $0.003763 | $0.003763 |
| 10,000 | $0.007526 | $0.007526 | $0.007526 |
| 25,000 | $0.01882 | $0.01882 | $0.01882 |
| 50,000 | $0.03763 | $0.03763 | $0.03763 |
| 100,000 | $0.07526 | $0.07526 | $0.07526 |
| 250,000 | $0.18815 | $0.18815 | $0.18815 |
| 500,000 | $0.3763 | $0.3763 | $0.3763 |
| 1,000,000 | $0.7526 | $0.7526 | $0.7526 |
| 5,000,000 | $3.763 | $3.763 | $3.763 |
| 10,000,000 | $7.526 | $7.526 | $7.526 |
| 50,000,000 | $37.63 | $37.63 | $37.63 |
| 100,000,000 | $75.26 | $75.26 | $75.26 |
| 500,000,000 | $376.30 | $376.30 | $376.30 |
| 1,000,000,000 | $752.60 | $752.60 | $752.60 |
| 10,000,000,000 | $7,526.00 | $7,526.00 | $7,526.00 |
| 100,000,000,000 | $75,260.00 | $75,260.00 | $75,260.00 |
Tokens for a fixed budget
How far a fixed spend goes on Llama 3.3 70B (watsonx.ai) at these list rates.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 13,287 | 13,287 | 13,287 |
| $0.1 | 132,873 | 132,873 | 132,873 |
| $0.5 | 664,364 | 664,364 | 664,364 |
| $1.00 | 1.33M | 1.33M | 1.33M |
| $5.00 | 6.64M | 6.64M | 6.64M |
| $10.00 | 13.3M | 13.3M | 13.3M |
| $20.00 | 26.6M | 26.6M | 26.6M |
| $25.00 | 33.2M | 33.2M | 33.2M |
| $30.00 | 39.9M | 39.9M | 39.9M |
| $50.00 | 66.4M | 66.4M | 66.4M |
| $100.00 | 132.9M | 132.9M | 132.9M |
| $250.00 | 332.2M | 332.2M | 332.2M |
| $500.00 | 664.4M | 664.4M | 664.4M |
| $1,000.00 | 1.33B | 1.33B | 1.33B |
| $5,000.00 | 6.64B | 6.64B | 6.64B |
| $10,000.00 | 13.3B | 13.3B | 13.3B |
At these rates
- A workload with 70% input and 30% output costs $0.7526 per 1 million total tokens and $75.26 per 100 million.
- At the same token count, output costs 1 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 128,000-token context window with input alone would cost about $0.09633, before any output.
IBM lists a single pay-as-you-go figure for this model row; treat input and output at the same published rate and confirm the live table.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.