DeepInfra models

Low-cost serverless inference for open models with transparent per-token rates.

https://deepinfra.com/pricing

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

There are 9 published models here. On a 70/30 mix, combined list rates run from $0.0223 to $1.69 per 1 million tokens, with a median of $0.16. The largest listed context window in this group is 1,024,000 tokens. The most recent price check in this group is 2026-08-14.

Official DeepInfra pricing page

DeepInfra list rates

USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.

ModelInput / 1MOutput / 1MMixed 70/30ContextVerified
Mistral Nemo (DeepInfra) $0.019 $0.03 $0.0223 128,000
Llama 3.1 8B Turbo (DeepInfra) $0.02 $0.04 $0.026 128,000
DeepSeek V4 Flash (DeepInfra) $0.09 $0.18 $0.117 1,024,000
Qwen3 32B (DeepInfra) $0.08 $0.28 $0.14 40,000
Llama 4 Scout (DeepInfra) $0.1 $0.3 $0.16 320,000
Gemma 4 31B Turbo (DeepInfra) $0.09 $0.34 $0.165 256,000
Llama 3.3 70B Turbo (DeepInfra) $0.1 $0.32 $0.166 128,000
Llama 4 Maverick (DeepInfra) $0.2 $0.8 $0.38 1,024,000
DeepSeek V4 Pro (DeepInfra) $1.30 $2.60 $1.69 1,024,000

Questions about DeepInfra rates

Do these DeepInfra figures include discounts and tax?

No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.

When is the more expensive DeepInfra tier worth the extra?

DeepSeek V4 Pro (DeepInfra) sits near $1.69 per 1M tokens on a 70/30 mix, versus $0.0223 on Mistral Nemo (DeepInfra). Use the high tier when retries or long reasoning actually fail on the cheap one.