Groq
Llama 3.1 8B Instant (Groq) token cost
Ultra-cheap Groq 8B tier for classification, routing, and lightweight agents at extreme speed.

Published list rates
Input
$0.05
per 1M tokens
Output
$0.08
per 1M tokens
Context window
131,072
tokens
Pricing source: Groq pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
Llama 3.1 8B Instant (Groq) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.00000005 | $0.00000008 | $0.000000059 |
| 10 | $0.0000005 | $0.0000008 | $0.00000059 |
| 100 | $0.000005 | $0.000008 | $0.000006 |
| 200 | $0.00001 | $0.000016 | $0.000012 |
| 300 | $0.000015 | $0.000024 | $0.000018 |
| 400 | $0.00002 | $0.000032 | $0.000024 |
| 500 | $0.000025 | $0.00004 | $0.00003 |
| 1,000 | $0.00005 | $0.00008 | $0.000059 |
| 2,000 | $0.0001 | $0.00016 | $0.000118 |
| 5,000 | $0.00025 | $0.0004 | $0.000295 |
| 10,000 | $0.0005 | $0.0008 | $0.00059 |
| 25,000 | $0.00125 | $0.002 | $0.001475 |
| 50,000 | $0.0025 | $0.004 | $0.00295 |
| 100,000 | $0.005 | $0.008 | $0.0059 |
| 250,000 | $0.0125 | $0.02 | $0.01475 |
| 500,000 | $0.025 | $0.04 | $0.0295 |
| 1,000,000 | $0.05 | $0.08 | $0.059 |
| 5,000,000 | $0.25 | $0.4 | $0.295 |
| 10,000,000 | $0.5 | $0.8 | $0.59 |
| 50,000,000 | $2.50 | $4.00 | $2.95 |
| 100,000,000 | $5.00 | $8.00 | $5.90 |
| 500,000,000 | $25.00 | $40.00 | $29.50 |
| 1,000,000,000 | $50.00 | $80.00 | $59.00 |
| 10,000,000,000 | $500.00 | $800.00 | $590.00 |
| 100,000,000,000 | $5,000.00 | $8,000.00 | $5,900.00 |
Tokens for a fixed budget
Tokens Llama 3.1 8B Instant (Groq) can process for a set budget — send-only, reply-only, and a 70/30 mix.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 200,000 | 125,000 | 169,492 |
| $0.1 | 2M | 1.25M | 1.69M |
| $0.5 | 10M | 6.25M | 8.47M |
| $1.00 | 20M | 12.5M | 16.9M |
| $5.00 | 100M | 62.5M | 84.7M |
| $10.00 | 200M | 125M | 169.5M |
| $20.00 | 400M | 250M | 339M |
| $25.00 | 500M | 312.5M | 423.7M |
| $30.00 | 600M | 375M | 508.5M |
| $50.00 | 1B | 625M | 847.5M |
| $100.00 | 2B | 1.25B | 1.69B |
| $250.00 | 5B | 3.13B | 4.24B |
| $500.00 | 10B | 6.25B | 8.47B |
| $1,000.00 | 20B | 12.5B | 16.9B |
| $5,000.00 | 100B | 62.5B | 84.7B |
| $10,000.00 | 200B | 125B | 169.5B |
At these rates
- A workload with 70% input and 30% output costs $0.059 per 1 million total tokens and $5.90 per 100 million.
- At the same token count, output costs 1.6 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 131,072-token context window with input alone would cost about $0.006554, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.