Cerebras

Gemma 4 31B (Cerebras) token cost

Google DeepMind Gemma 4 31B on Cerebras Developer tier for high-speed open inference.

Published list rates

Input

$0.99

per 1M tokens

Output

$1.49

per 1M tokens

Context window

128,000

tokens

Pricing source: Cerebras pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Read down to see Gemma 4 31B (Cerebras) get more expensive as tokens grow. The last column is a typical mix: 70% input, 30% output.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000099 $0.000001 $0.000001
10 $0.00001 $0.000015 $0.000011
100 $0.000099 $0.000149 $0.000114
200 $0.000198 $0.000298 $0.000228
300 $0.000297 $0.000447 $0.000342
400 $0.000396 $0.000596 $0.000456
500 $0.000495 $0.000745 $0.00057
1,000 $0.00099 $0.00149 $0.00114
2,000 $0.00198 $0.00298 $0.00228
5,000 $0.00495 $0.00745 $0.0057
10,000 $0.0099 $0.0149 $0.0114
25,000 $0.02475 $0.03725 $0.0285
50,000 $0.0495 $0.0745 $0.057
100,000 $0.099 $0.149 $0.114
250,000 $0.2475 $0.3725 $0.285
500,000 $0.495 $0.745 $0.57
1,000,000 $0.99 $1.49 $1.14
5,000,000 $4.95 $7.45 $5.70
10,000,000 $9.90 $14.90 $11.40
50,000,000 $49.50 $74.50 $57.00
100,000,000 $99.00 $149.00 $114.00
500,000,000 $495.00 $745.00 $570.00
1,000,000,000 $990.00 $1,490.00 $1,140.00
10,000,000,000 $9,900.00 $14,900.00 $11,400.00
100,000,000,000 $99,000.00 $149,000.00 $114,000.00

Tokens for a fixed budget

If you cap spend, this is roughly how many Gemma 4 31B (Cerebras) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 10,101 6,711 8,772
$0.1 101,010 67,114 87,719
$0.5 505,051 335,570 438,596
$1.00 1.01M 671,141 877,193
$5.00 5.05M 3.36M 4.39M
$10.00 10.1M 6.71M 8.77M
$20.00 20.2M 13.4M 17.5M
$25.00 25.3M 16.8M 21.9M
$30.00 30.3M 20.1M 26.3M
$50.00 50.5M 33.6M 43.9M
$100.00 101M 67.1M 87.7M
$250.00 252.5M 167.8M 219.3M
$500.00 505.1M 335.6M 438.6M
$1,000.00 1.01B 671.1M 877.2M
$5,000.00 5.05B 3.36B 4.39B
$10,000.00 10.1B 6.71B 8.77B

At these rates

  • A workload with 70% input and 30% output costs $1.14 per 1 million total tokens and $114.00 per 100 million.
  • At the same token count, output costs 1.51 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 128,000-token context window with input alone would cost about $0.12672, before any output.

Cerebras

Wafer-scale inference with extreme tokens-per-second on open models.

All Cerebras models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models