Fireworks models
Serverless open-model inference with per-token Standard and Priority paths.


Kimi K3 (Fireworks)
$3.00 / $15.00 per 1M tokens

DeepSeek V4 Pro (Fireworks)
$1.74 / $3.48 per 1M tokens

DeepSeek V4 Flash (Fireworks)
$0.14 / $0.28 per 1M tokens

GLM 5.2 (Fireworks)
$1.40 / $4.40 per 1M tokens

MiniMax M3 (Fireworks)
$0.3 / $1.20 per 1M tokens

GPT OSS 120B (Fireworks)
$0.15 / $0.6 per 1M tokens

GPT OSS 20B (Fireworks)
$0.07 / $0.3 per 1M tokens

Kimi K2.6 (Fireworks)
$0.95 / $4.00 per 1M tokens
Fireworks
Serverless open-model inference with per-token Standard and Priority paths.
There are 8 published models here. On a 70/30 mix, combined list rates run from $0.139 to $6.60 per 1 million tokens, with a median of $1.2175. The largest listed context window in this group is 1,000,000 tokens. The most recent price check in this group is 2026-08-14.
Official Fireworks pricing page
Fireworks list rates
USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.
| Model | Input / 1M | Output / 1M | Mixed 70/30 | Context | Verified |
|---|---|---|---|---|---|
| GPT OSS 20B (Fireworks) | $0.07 | $0.3 | $0.139 | 131,072 | |
| DeepSeek V4 Flash (Fireworks) | $0.14 | $0.28 | $0.182 | 256,000 | |
| GPT OSS 120B (Fireworks) | $0.15 | $0.6 | $0.285 | 131,072 | |
| MiniMax M3 (Fireworks) | $0.3 | $1.20 | $0.57 | 1,000,000 | |
| Kimi K2.6 (Fireworks) | $0.95 | $4.00 | $1.865 | 256,000 | |
| DeepSeek V4 Pro (Fireworks) | $1.74 | $3.48 | $2.262 | 512,000 | |
| GLM 5.2 (Fireworks) | $1.40 | $4.40 | $2.30 | 200,000 | |
| Kimi K3 (Fireworks) | $3.00 | $15.00 | $6.60 | 256,000 |
Questions about Fireworks rates
Do these Fireworks figures include discounts and tax?
No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.
When is the more expensive Fireworks tier worth the extra?
Kimi K3 (Fireworks) sits near $6.60 per 1M tokens on a 70/30 mix, versus $0.139 on GPT OSS 20B (Fireworks). Use the high tier when retries or long reasoning actually fail on the cheap one.