Cerebras models

Wafer-scale inference with extreme tokens-per-second on open models.

https://www.cerebras.ai/pricing

Cerebras

Wafer-scale inference with extreme tokens-per-second on open models.

There are 3 published models here. On a 70/30 mix, combined list rates run from $0.47 to $2.40 per 1 million tokens, with a median of $1.14. The largest listed context window in this group is 131,072 tokens. The most recent price check in this group is 2026-08-14.

Official Cerebras pricing page

Cerebras list rates

USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.

ModelInput / 1MOutput / 1MMixed 70/30ContextVerified
GPT OSS 120B (Cerebras) $0.35 $0.75 $0.47 131,072
Gemma 4 31B (Cerebras) $0.99 $1.49 $1.14 128,000
GLM 4.7 (Cerebras) $2.25 $2.75 $2.40 128,000

Questions about Cerebras rates

Do these Cerebras figures include discounts and tax?

No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.

When is the more expensive Cerebras tier worth the extra?

GLM 4.7 (Cerebras) sits near $2.40 per 1M tokens on a 70/30 mix, versus $0.47 on GPT OSS 120B (Cerebras). Use the high tier when retries or long reasoning actually fail on the cheap one.