Cerebras models
Wafer-scale inference with extreme tokens-per-second on open models.


GPT OSS 120B (Cerebras)
$0.35 / $0.75 per 1M tokens

Gemma 4 31B (Cerebras)
$0.99 / $1.49 per 1M tokens

GLM 4.7 (Cerebras)
$2.25 / $2.75 per 1M tokens
Cerebras
Wafer-scale inference with extreme tokens-per-second on open models.
There are 3 published models here. On a 70/30 mix, combined list rates run from $0.47 to $2.40 per 1 million tokens, with a median of $1.14. The largest listed context window in this group is 131,072 tokens. The most recent price check in this group is 2026-08-14.
Official Cerebras pricing page
Cerebras list rates
USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.
| Model | Input / 1M | Output / 1M | Mixed 70/30 | Context | Verified |
|---|---|---|---|---|---|
| GPT OSS 120B (Cerebras) | $0.35 | $0.75 | $0.47 | 131,072 | |
| Gemma 4 31B (Cerebras) | $0.99 | $1.49 | $1.14 | 128,000 | |
| GLM 4.7 (Cerebras) | $2.25 | $2.75 | $2.40 | 128,000 |
Questions about Cerebras rates
Do these Cerebras figures include discounts and tax?
No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.
When is the more expensive Cerebras tier worth the extra?
GLM 4.7 (Cerebras) sits near $2.40 per 1M tokens on a 70/30 mix, versus $0.47 on GPT OSS 120B (Cerebras). Use the high tier when retries or long reasoning actually fail on the cheap one.