Cerebras
Gemma 4 31B (Cerebras) token cost
Google DeepMind Gemma 4 31B on Cerebras Developer tier for high-speed open inference.

Published list rates
Input
$0.99
per 1M tokens
Output
$1.49
per 1M tokens
Context window
128,000
tokens
Pricing source: Cerebras pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
Read down to see Gemma 4 31B (Cerebras) get more expensive as tokens grow. The last column is a typical mix: 70% input, 30% output.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.00000099 | $0.000001 | $0.000001 |
| 10 | $0.00001 | $0.000015 | $0.000011 |
| 100 | $0.000099 | $0.000149 | $0.000114 |
| 200 | $0.000198 | $0.000298 | $0.000228 |
| 300 | $0.000297 | $0.000447 | $0.000342 |
| 400 | $0.000396 | $0.000596 | $0.000456 |
| 500 | $0.000495 | $0.000745 | $0.00057 |
| 1,000 | $0.00099 | $0.00149 | $0.00114 |
| 2,000 | $0.00198 | $0.00298 | $0.00228 |
| 5,000 | $0.00495 | $0.00745 | $0.0057 |
| 10,000 | $0.0099 | $0.0149 | $0.0114 |
| 25,000 | $0.02475 | $0.03725 | $0.0285 |
| 50,000 | $0.0495 | $0.0745 | $0.057 |
| 100,000 | $0.099 | $0.149 | $0.114 |
| 250,000 | $0.2475 | $0.3725 | $0.285 |
| 500,000 | $0.495 | $0.745 | $0.57 |
| 1,000,000 | $0.99 | $1.49 | $1.14 |
| 5,000,000 | $4.95 | $7.45 | $5.70 |
| 10,000,000 | $9.90 | $14.90 | $11.40 |
| 50,000,000 | $49.50 | $74.50 | $57.00 |
| 100,000,000 | $99.00 | $149.00 | $114.00 |
| 500,000,000 | $495.00 | $745.00 | $570.00 |
| 1,000,000,000 | $990.00 | $1,490.00 | $1,140.00 |
| 10,000,000,000 | $9,900.00 | $14,900.00 | $11,400.00 |
| 100,000,000,000 | $99,000.00 | $149,000.00 | $114,000.00 |
Tokens for a fixed budget
If you cap spend, this is roughly how many Gemma 4 31B (Cerebras) tokens that money buys.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 10,101 | 6,711 | 8,772 |
| $0.1 | 101,010 | 67,114 | 87,719 |
| $0.5 | 505,051 | 335,570 | 438,596 |
| $1.00 | 1.01M | 671,141 | 877,193 |
| $5.00 | 5.05M | 3.36M | 4.39M |
| $10.00 | 10.1M | 6.71M | 8.77M |
| $20.00 | 20.2M | 13.4M | 17.5M |
| $25.00 | 25.3M | 16.8M | 21.9M |
| $30.00 | 30.3M | 20.1M | 26.3M |
| $50.00 | 50.5M | 33.6M | 43.9M |
| $100.00 | 101M | 67.1M | 87.7M |
| $250.00 | 252.5M | 167.8M | 219.3M |
| $500.00 | 505.1M | 335.6M | 438.6M |
| $1,000.00 | 1.01B | 671.1M | 877.2M |
| $5,000.00 | 5.05B | 3.36B | 4.39B |
| $10,000.00 | 10.1B | 6.71B | 8.77B |
At these rates
- A workload with 70% input and 30% output costs $1.14 per 1 million total tokens and $114.00 per 100 million.
- At the same token count, output costs 1.51 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 128,000-token context window with input alone would cost about $0.12672, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.