Cerebras
GLM 4.7 (Cerebras) token cost
Z.AI GLM 4.7 preview rates on Cerebras. Confirm deprecation dates on the public pricing table.

Published list rates
Input
$2.25
per 1M tokens
Output
$2.75
per 1M tokens
Context window
128,000
tokens
Pricing source: Cerebras pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
What GLM 4.7 (Cerebras) costs if you only send tokens, only get a reply, or do both (70% in / 30% out). Rows start small and go up to huge traffic.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.000002 | $0.000003 | $0.000002 |
| 10 | $0.000023 | $0.000028 | $0.000024 |
| 100 | $0.000225 | $0.000275 | $0.00024 |
| 200 | $0.00045 | $0.00055 | $0.00048 |
| 300 | $0.000675 | $0.000825 | $0.00072 |
| 400 | $0.0009 | $0.0011 | $0.00096 |
| 500 | $0.001125 | $0.001375 | $0.0012 |
| 1,000 | $0.00225 | $0.00275 | $0.0024 |
| 2,000 | $0.0045 | $0.0055 | $0.0048 |
| 5,000 | $0.01125 | $0.01375 | $0.012 |
| 10,000 | $0.0225 | $0.0275 | $0.024 |
| 25,000 | $0.05625 | $0.06875 | $0.06 |
| 50,000 | $0.1125 | $0.1375 | $0.12 |
| 100,000 | $0.225 | $0.275 | $0.24 |
| 250,000 | $0.5625 | $0.6875 | $0.6 |
| 500,000 | $1.125 | $1.375 | $1.20 |
| 1,000,000 | $2.25 | $2.75 | $2.40 |
| 5,000,000 | $11.25 | $13.75 | $12.00 |
| 10,000,000 | $22.50 | $27.50 | $24.00 |
| 50,000,000 | $112.50 | $137.50 | $120.00 |
| 100,000,000 | $225.00 | $275.00 | $240.00 |
| 500,000,000 | $1,125.00 | $1,375.00 | $1,200.00 |
| 1,000,000,000 | $2,250.00 | $2,750.00 | $2,400.00 |
| 10,000,000,000 | $22,500.00 | $27,500.00 | $24,000.00 |
| 100,000,000,000 | $225,000.00 | $275,000.00 | $240,000.00 |
Tokens for a fixed budget
How far a fixed spend goes on GLM 4.7 (Cerebras) at these list rates.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 4,444 | 3,636 | 4,167 |
| $0.1 | 44,444 | 36,364 | 41,667 |
| $0.5 | 222,222 | 181,818 | 208,333 |
| $1.00 | 444,444 | 363,636 | 416,667 |
| $5.00 | 2.22M | 1.82M | 2.08M |
| $10.00 | 4.44M | 3.64M | 4.17M |
| $20.00 | 8.89M | 7.27M | 8.33M |
| $25.00 | 11.1M | 9.09M | 10.4M |
| $30.00 | 13.3M | 10.9M | 12.5M |
| $50.00 | 22.2M | 18.2M | 20.8M |
| $100.00 | 44.4M | 36.4M | 41.7M |
| $250.00 | 111.1M | 90.9M | 104.2M |
| $500.00 | 222.2M | 181.8M | 208.3M |
| $1,000.00 | 444.4M | 363.6M | 416.7M |
| $5,000.00 | 2.22B | 1.82B | 2.08B |
| $10,000.00 | 4.44B | 3.64B | 4.17B |
At these rates
- A workload with 70% input and 30% output costs $2.40 per 1 million total tokens and $240.00 per 100 million.
- At the same token count, output costs 1.22 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 128,000-token context window with input alone would cost about $0.288, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.