Novita AI
Llama 3.3 70B Instruct (Novita) token cost
Llama 3.3 70B on Novita for general chat — compare with DeepInfra and Groq unit rates.

Published list rates
Input
$0.135
per 1M tokens
Output
$0.4
per 1M tokens
Context window
128,000
tokens
Pricing source: Novita AI pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
Three views of Llama 3.3 70B Instruct (Novita): sending only, receiving only, and a 70/30 mix. Same list rates, bigger rows.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.000000135 | $0.0000004 | $0.0000002145 |
| 10 | $0.000001 | $0.000004 | $0.000002 |
| 100 | $0.000014 | $0.00004 | $0.000021 |
| 200 | $0.000027 | $0.00008 | $0.000043 |
| 300 | $0.000041 | $0.00012 | $0.000064 |
| 400 | $0.000054 | $0.00016 | $0.000086 |
| 500 | $0.000068 | $0.0002 | $0.000107 |
| 1,000 | $0.000135 | $0.0004 | $0.000215 |
| 2,000 | $0.00027 | $0.0008 | $0.000429 |
| 5,000 | $0.000675 | $0.002 | $0.001073 |
| 10,000 | $0.00135 | $0.004 | $0.002145 |
| 25,000 | $0.003375 | $0.01 | $0.005363 |
| 50,000 | $0.00675 | $0.02 | $0.01073 |
| 100,000 | $0.0135 | $0.04 | $0.02145 |
| 250,000 | $0.03375 | $0.1 | $0.05363 |
| 500,000 | $0.0675 | $0.2 | $0.10725 |
| 1,000,000 | $0.135 | $0.4 | $0.2145 |
| 5,000,000 | $0.675 | $2.00 | $1.0725 |
| 10,000,000 | $1.35 | $4.00 | $2.145 |
| 50,000,000 | $6.75 | $20.00 | $10.725 |
| 100,000,000 | $13.50 | $40.00 | $21.45 |
| 500,000,000 | $67.50 | $200.00 | $107.25 |
| 1,000,000,000 | $135.00 | $400.00 | $214.50 |
| 10,000,000,000 | $1,350.00 | $4,000.00 | $2,145.00 |
| 100,000,000,000 | $13,500.00 | $40,000.00 | $21,450.00 |
Tokens for a fixed budget
Tokens Llama 3.3 70B Instruct (Novita) can process for a set budget — send-only, reply-only, and a 70/30 mix.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 74,074 | 25,000 | 46,620 |
| $0.1 | 740,741 | 250,000 | 466,200 |
| $0.5 | 3.7M | 1.25M | 2.33M |
| $1.00 | 7.41M | 2.5M | 4.66M |
| $5.00 | 37M | 12.5M | 23.3M |
| $10.00 | 74.1M | 25M | 46.6M |
| $20.00 | 148.1M | 50M | 93.2M |
| $25.00 | 185.2M | 62.5M | 116.6M |
| $30.00 | 222.2M | 75M | 139.9M |
| $50.00 | 370.4M | 125M | 233.1M |
| $100.00 | 740.7M | 250M | 466.2M |
| $250.00 | 1.85B | 625M | 1.17B |
| $500.00 | 3.7B | 1.25B | 2.33B |
| $1,000.00 | 7.41B | 2.5B | 4.66B |
| $5,000.00 | 37B | 12.5B | 23.3B |
| $10,000.00 | 74.1B | 25B | 46.6B |
At these rates
- A workload with 70% input and 30% output costs $0.2145 per 1 million total tokens and $21.45 per 100 million.
- At the same token count, output costs 2.96 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 128,000-token context window with input alone would cost about $0.01728, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.