Z.AI
GLM-5-Turbo token cost
Faster GLM-5-Turbo first-party tier when latency matters more than base GLM-5 unit cost.

Published list rates
Input
$1.20
per 1M tokens
Output
$4.00
per 1M tokens
Cached input
$0.24
per 1M tokens
Context window
200,000
tokens
Pricing source: Z.AI pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
What GLM-5-Turbo costs if you only send tokens, only get a reply, or do both (70% in / 30% out). Rows start small and go up to huge traffic.
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.000001 | $0.000004 | $0.000002 |
| 10 | $0.000012 | $0.00004 | $0.00002 |
| 100 | $0.00012 | $0.0004 | $0.000204 |
| 200 | $0.00024 | $0.0008 | $0.000408 |
| 300 | $0.00036 | $0.0012 | $0.000612 |
| 400 | $0.00048 | $0.0016 | $0.000816 |
| 500 | $0.0006 | $0.002 | $0.00102 |
| 1,000 | $0.0012 | $0.004 | $0.00204 |
| 2,000 | $0.0024 | $0.008 | $0.00408 |
| 5,000 | $0.006 | $0.02 | $0.0102 |
| 10,000 | $0.012 | $0.04 | $0.0204 |
| 25,000 | $0.03 | $0.1 | $0.051 |
| 50,000 | $0.06 | $0.2 | $0.102 |
| 100,000 | $0.12 | $0.4 | $0.204 |
| 250,000 | $0.3 | $1.00 | $0.51 |
| 500,000 | $0.6 | $2.00 | $1.02 |
| 1,000,000 | $1.20 | $4.00 | $2.04 |
| 5,000,000 | $6.00 | $20.00 | $10.20 |
| 10,000,000 | $12.00 | $40.00 | $20.40 |
| 50,000,000 | $60.00 | $200.00 | $102.00 |
| 100,000,000 | $120.00 | $400.00 | $204.00 |
| 500,000,000 | $600.00 | $2,000.00 | $1,020.00 |
| 1,000,000,000 | $1,200.00 | $4,000.00 | $2,040.00 |
| 10,000,000,000 | $12,000.00 | $40,000.00 | $20,400.00 |
| 100,000,000,000 | $120,000.00 | $400,000.00 | $204,000.00 |
Tokens for a fixed budget
If you cap spend, this is roughly how many GLM-5-Turbo tokens that money buys.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 8,333 | 2,500 | 4,902 |
| $0.1 | 83,333 | 25,000 | 49,020 |
| $0.5 | 416,667 | 125,000 | 245,098 |
| $1.00 | 833,333 | 250,000 | 490,196 |
| $5.00 | 4.17M | 1.25M | 2.45M |
| $10.00 | 8.33M | 2.5M | 4.9M |
| $20.00 | 16.7M | 5M | 9.8M |
| $25.00 | 20.8M | 6.25M | 12.3M |
| $30.00 | 25M | 7.5M | 14.7M |
| $50.00 | 41.7M | 12.5M | 24.5M |
| $100.00 | 83.3M | 25M | 49M |
| $250.00 | 208.3M | 62.5M | 122.5M |
| $500.00 | 416.7M | 125M | 245.1M |
| $1,000.00 | 833.3M | 250M | 490.2M |
| $5,000.00 | 4.17B | 1.25B | 2.45B |
| $10,000.00 | 8.33B | 2.5B | 4.9B |
At these rates
- A workload with 70% input and 30% output costs $2.04 per 1 million total tokens and $204.00 per 100 million.
- At the same token count, output costs 3.33 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 200,000-token context window with input alone would cost about $0.24, before any output.
- Cached input is $0.24 per million, 80% below standard input when the workload meets the provider’s cache rules.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.