Cerebras

GPT OSS 120B (Cerebras) token cost

GPT OSS 120B on Cerebras wafer-scale inference — compare host rates with Groq, Fireworks, and Together.

Published list rates

Input

$0.35

per 1M tokens

Output

$0.75

per 1M tokens

Context window

131,072

tokens

Pricing source: Cerebras pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

What GPT OSS 120B (Cerebras) costs if you only send tokens, only get a reply, or do both (70% in / 30% out). Rows start small and go up to huge traffic.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000035 $0.00000075 $0.00000047
10 $0.000004 $0.000008 $0.000005
100 $0.000035 $0.000075 $0.000047
200 $0.00007 $0.00015 $0.000094
300 $0.000105 $0.000225 $0.000141
400 $0.00014 $0.0003 $0.000188
500 $0.000175 $0.000375 $0.000235
1,000 $0.00035 $0.00075 $0.00047
2,000 $0.0007 $0.0015 $0.00094
5,000 $0.00175 $0.00375 $0.00235
10,000 $0.0035 $0.0075 $0.0047
25,000 $0.00875 $0.01875 $0.01175
50,000 $0.0175 $0.0375 $0.0235
100,000 $0.035 $0.075 $0.047
250,000 $0.0875 $0.1875 $0.1175
500,000 $0.175 $0.375 $0.235
1,000,000 $0.35 $0.75 $0.47
5,000,000 $1.75 $3.75 $2.35
10,000,000 $3.50 $7.50 $4.70
50,000,000 $17.50 $37.50 $23.50
100,000,000 $35.00 $75.00 $47.00
500,000,000 $175.00 $375.00 $235.00
1,000,000,000 $350.00 $750.00 $470.00
10,000,000,000 $3,500.00 $7,500.00 $4,700.00
100,000,000,000 $35,000.00 $75,000.00 $47,000.00

Tokens for a fixed budget

If you cap spend, this is roughly how many GPT OSS 120B (Cerebras) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 28,571 13,333 21,277
$0.1 285,714 133,333 212,766
$0.5 1.43M 666,667 1.06M
$1.00 2.86M 1.33M 2.13M
$5.00 14.3M 6.67M 10.6M
$10.00 28.6M 13.3M 21.3M
$20.00 57.1M 26.7M 42.6M
$25.00 71.4M 33.3M 53.2M
$30.00 85.7M 40M 63.8M
$50.00 142.9M 66.7M 106.4M
$100.00 285.7M 133.3M 212.8M
$250.00 714.3M 333.3M 531.9M
$500.00 1.43B 666.7M 1.06B
$1,000.00 2.86B 1.33B 2.13B
$5,000.00 14.3B 6.67B 10.6B
$10,000.00 28.6B 13.3B 21.3B

At these rates

  • A workload with 70% input and 30% output costs $0.47 per 1 million total tokens and $47.00 per 100 million.
  • At the same token count, output costs 2.14 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 131,072-token context window with input alone would cost about $0.04588, before any output.

Cerebras

Wafer-scale inference with extreme tokens-per-second on open models.

All Cerebras models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models