Baseten

GLM-5.2 Fast (Baseten) token cost

Faster GLM-5.2 variant on Baseten when latency matters more than token price.

Published list rates

Input

$2.10

per 1M tokens

Output

$6.60

per 1M tokens

Cached input

$0.21

per 1M tokens

Context window

200,000

tokens

Pricing source: Baseten Model APIs pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Read down to see GLM-5.2 Fast (Baseten) get more expensive as tokens grow. The last column is a typical mix: 70% input, 30% output.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.000002 $0.000007 $0.000003
10 $0.000021 $0.000066 $0.000035
100 $0.00021 $0.00066 $0.000345
200 $0.00042 $0.00132 $0.00069
300 $0.00063 $0.00198 $0.001035
400 $0.00084 $0.00264 $0.00138
500 $0.00105 $0.0033 $0.001725
1,000 $0.0021 $0.0066 $0.00345
2,000 $0.0042 $0.0132 $0.0069
5,000 $0.0105 $0.033 $0.01725
10,000 $0.021 $0.066 $0.0345
25,000 $0.0525 $0.165 $0.08625
50,000 $0.105 $0.33 $0.1725
100,000 $0.21 $0.66 $0.345
250,000 $0.525 $1.65 $0.8625
500,000 $1.05 $3.30 $1.725
1,000,000 $2.10 $6.60 $3.45
5,000,000 $10.50 $33.00 $17.25
10,000,000 $21.00 $66.00 $34.50
50,000,000 $105.00 $330.00 $172.50
100,000,000 $210.00 $660.00 $345.00
500,000,000 $1,050.00 $3,300.00 $1,725.00
1,000,000,000 $2,100.00 $6,600.00 $3,450.00
10,000,000,000 $21,000.00 $66,000.00 $34,500.00
100,000,000,000 $210,000.00 $660,000.00 $345,000.00

Tokens for a fixed budget

If you cap spend, this is roughly how many GLM-5.2 Fast (Baseten) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 4,762 1,515 2,899
$0.1 47,619 15,152 28,986
$0.5 238,095 75,758 144,928
$1.00 476,190 151,515 289,855
$5.00 2.38M 757,576 1.45M
$10.00 4.76M 1.52M 2.9M
$20.00 9.52M 3.03M 5.8M
$25.00 11.9M 3.79M 7.25M
$30.00 14.3M 4.55M 8.7M
$50.00 23.8M 7.58M 14.5M
$100.00 47.6M 15.2M 29M
$250.00 119M 37.9M 72.5M
$500.00 238.1M 75.8M 144.9M
$1,000.00 476.2M 151.5M 289.9M
$5,000.00 2.38B 757.6M 1.45B
$10,000.00 4.76B 1.52B 2.9B

At these rates

  • A workload with 70% input and 30% output costs $3.45 per 1 million total tokens and $345.00 per 100 million.
  • At the same token count, output costs 3.14 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 200,000-token context window with input alone would cost about $0.42, before any output.
  • Cached input is $0.21 per million, 90% below standard input when the workload meets the provider’s cache rules.

Baseten

Model APIs and dedicated GPU inference with public per-token list rates.

All Baseten models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models