Z.AI
GLM-5 token cost
Z.AI GLM-5 first-party rates with cached input for coding and agentic engineering.

Published list rates
Input
$1.00
per 1M tokens
Output
$3.20
per 1M tokens
Cached input
$0.2
per 1M tokens
Context window
200,000
tokens
Pricing source: Z.AI pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
GLM-5 at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.000001 | $0.000003 | $0.000002 |
| 10 | $0.00001 | $0.000032 | $0.000017 |
| 100 | $0.0001 | $0.00032 | $0.000166 |
| 200 | $0.0002 | $0.00064 | $0.000332 |
| 300 | $0.0003 | $0.00096 | $0.000498 |
| 400 | $0.0004 | $0.00128 | $0.000664 |
| 500 | $0.0005 | $0.0016 | $0.00083 |
| 1,000 | $0.001 | $0.0032 | $0.00166 |
| 2,000 | $0.002 | $0.0064 | $0.00332 |
| 5,000 | $0.005 | $0.016 | $0.0083 |
| 10,000 | $0.01 | $0.032 | $0.0166 |
| 25,000 | $0.025 | $0.08 | $0.0415 |
| 50,000 | $0.05 | $0.16 | $0.083 |
| 100,000 | $0.1 | $0.32 | $0.166 |
| 250,000 | $0.25 | $0.8 | $0.415 |
| 500,000 | $0.5 | $1.60 | $0.83 |
| 1,000,000 | $1.00 | $3.20 | $1.66 |
| 5,000,000 | $5.00 | $16.00 | $8.30 |
| 10,000,000 | $10.00 | $32.00 | $16.60 |
| 50,000,000 | $50.00 | $160.00 | $83.00 |
| 100,000,000 | $100.00 | $320.00 | $166.00 |
| 500,000,000 | $500.00 | $1,600.00 | $830.00 |
| 1,000,000,000 | $1,000.00 | $3,200.00 | $1,660.00 |
| 10,000,000,000 | $10,000.00 | $32,000.00 | $16,600.00 |
| 100,000,000,000 | $100,000.00 | $320,000.00 | $166,000.00 |
Tokens for a fixed budget
If you cap spend, this is roughly how many GLM-5 tokens that money buys.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 10,000 | 3,125 | 6,024 |
| $0.1 | 100,000 | 31,250 | 60,241 |
| $0.5 | 500,000 | 156,250 | 301,205 |
| $1.00 | 1M | 312,500 | 602,410 |
| $5.00 | 5M | 1.56M | 3.01M |
| $10.00 | 10M | 3.13M | 6.02M |
| $20.00 | 20M | 6.25M | 12M |
| $25.00 | 25M | 7.81M | 15.1M |
| $30.00 | 30M | 9.38M | 18.1M |
| $50.00 | 50M | 15.6M | 30.1M |
| $100.00 | 100M | 31.3M | 60.2M |
| $250.00 | 250M | 78.1M | 150.6M |
| $500.00 | 500M | 156.3M | 301.2M |
| $1,000.00 | 1B | 312.5M | 602.4M |
| $5,000.00 | 5B | 1.56B | 3.01B |
| $10,000.00 | 10B | 3.13B | 6.02B |
At these rates
- A workload with 70% input and 30% output costs $1.66 per 1 million total tokens and $166.00 per 100 million.
- At the same token count, output costs 3.2 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 200,000-token context window with input alone would cost about $0.2, before any output.
- Cached input is $0.2 per million, 80% below standard input when the workload meets the provider’s cache rules.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.