SiliconFlow models
Serverless open-model inference with public per-token rates for DeepSeek, Qwen, GLM, and more.

DeepSeek V4 Flash (SiliconFlow)
$0.13 / $0.28 per 1M tokens
DeepSeek V3.2 (SiliconFlow)
$0.27 / $0.42 per 1M tokens
Qwen3.5 9B (SiliconFlow)
$0.1 / $0.15 per 1M tokens
GLM-5.2 (SiliconFlow)
$1.302 / $4.092 per 1M tokens
Kimi K3 (SiliconFlow)
$3.00 / $15.00 per 1M tokens
gpt-oss-120b (SiliconFlow)
$0.05 / $0.45 per 1M tokens
MiniMax M3 (SiliconFlow)
$0.3 / $1.20 per 1M tokens
SiliconFlow
Serverless open-model inference with public per-token rates for DeepSeek, Qwen, GLM, and more.
There are 7 published models here. On a 70/30 mix, combined list rates run from $0.115 to $6.60 per 1 million tokens, with a median of $0.315. The largest listed context window in this group is 1,049,000 tokens. The most recent price check in this group is 2026-08-14.
Official SiliconFlow pricing page
SiliconFlow list rates
USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.
| Model | Input / 1M | Output / 1M | Mixed 70/30 | Context | Verified |
|---|---|---|---|---|---|
| Qwen3.5 9B (SiliconFlow) | $0.1 | $0.15 | $0.115 | 262,000 | |
| gpt-oss-120b (SiliconFlow) | $0.05 | $0.45 | $0.17 | 131,000 | |
| DeepSeek V4 Flash (SiliconFlow) | $0.13 | $0.28 | $0.175 | 1,049,000 | |
| DeepSeek V3.2 (SiliconFlow) | $0.27 | $0.42 | $0.315 | 164,000 | |
| MiniMax M3 (SiliconFlow) | $0.3 | $1.20 | $0.57 | 1,049,000 | |
| GLM-5.2 (SiliconFlow) | $1.302 | $4.092 | $2.139 | 1,049,000 | |
| Kimi K3 (SiliconFlow) | $3.00 | $15.00 | $6.60 | 1,049,000 |
Questions about SiliconFlow rates
Do these SiliconFlow figures include discounts and tax?
No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.
When is the more expensive SiliconFlow tier worth the extra?
Kimi K3 (SiliconFlow) sits near $6.60 per 1M tokens on a 70/30 mix, versus $0.115 on Qwen3.5 9B (SiliconFlow). Use the high tier when retries or long reasoning actually fail on the cheap one.