Groq
Llama 3.3 70B Versatile (Groq) token cost
Meta Llama 3.3 70B on Groq for high-throughput chat when latency and open weights both matter.

Published list rates
Input
$0.59
per 1M tokens
Output
$0.79
per 1M tokens
Context window
131,072
tokens
Pricing source: Groq pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
What Llama 3.3 70B Versatile (Groq) costs if you only send tokens, only get a reply, or do both (70% in / 30% out). Rows start small and go up to huge traffic.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.00000059 | $0.00000079 | $0.00000065 |
| 10 | $0.000006 | $0.000008 | $0.000007 |
| 100 | $0.000059 | $0.000079 | $0.000065 |
| 200 | $0.000118 | $0.000158 | $0.00013 |
| 300 | $0.000177 | $0.000237 | $0.000195 |
| 400 | $0.000236 | $0.000316 | $0.00026 |
| 500 | $0.000295 | $0.000395 | $0.000325 |
| 1,000 | $0.00059 | $0.00079 | $0.00065 |
| 2,000 | $0.00118 | $0.00158 | $0.0013 |
| 5,000 | $0.00295 | $0.00395 | $0.00325 |
| 10,000 | $0.0059 | $0.0079 | $0.0065 |
| 25,000 | $0.01475 | $0.01975 | $0.01625 |
| 50,000 | $0.0295 | $0.0395 | $0.0325 |
| 100,000 | $0.059 | $0.079 | $0.065 |
| 250,000 | $0.1475 | $0.1975 | $0.1625 |
| 500,000 | $0.295 | $0.395 | $0.325 |
| 1,000,000 | $0.59 | $0.79 | $0.65 |
| 5,000,000 | $2.95 | $3.95 | $3.25 |
| 10,000,000 | $5.90 | $7.90 | $6.50 |
| 50,000,000 | $29.50 | $39.50 | $32.50 |
| 100,000,000 | $59.00 | $79.00 | $65.00 |
| 500,000,000 | $295.00 | $395.00 | $325.00 |
| 1,000,000,000 | $590.00 | $790.00 | $650.00 |
| 10,000,000,000 | $5,900.00 | $7,900.00 | $6,500.00 |
| 100,000,000,000 | $59,000.00 | $79,000.00 | $65,000.00 |
Tokens for a fixed budget
Tokens Llama 3.3 70B Versatile (Groq) can process for a set budget — send-only, reply-only, and a 70/30 mix.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 16,949 | 12,658 | 15,385 |
| $0.1 | 169,492 | 126,582 | 153,846 |
| $0.5 | 847,458 | 632,911 | 769,231 |
| $1.00 | 1.69M | 1.27M | 1.54M |
| $5.00 | 8.47M | 6.33M | 7.69M |
| $10.00 | 16.9M | 12.7M | 15.4M |
| $20.00 | 33.9M | 25.3M | 30.8M |
| $25.00 | 42.4M | 31.6M | 38.5M |
| $30.00 | 50.8M | 38M | 46.2M |
| $50.00 | 84.7M | 63.3M | 76.9M |
| $100.00 | 169.5M | 126.6M | 153.8M |
| $250.00 | 423.7M | 316.5M | 384.6M |
| $500.00 | 847.5M | 632.9M | 769.2M |
| $1,000.00 | 1.69B | 1.27B | 1.54B |
| $5,000.00 | 8.47B | 6.33B | 7.69B |
| $10,000.00 | 16.9B | 12.7B | 15.4B |
At these rates
- A workload with 70% input and 30% output costs $0.65 per 1 million total tokens and $65.00 per 100 million.
- At the same token count, output costs 1.34 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 131,072-token context window with input alone would cost about $0.07733, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.