DeepInfra Models
Low-cost serverless inference for open models with transparent per-token rates.

Llama 3.3 70B Turbo (DeepInfra)
$0.1 / $0.32 per 1M tokens
Llama 3.1 8B Turbo (DeepInfra)
$0.02 / $0.04 per 1M tokens
DeepSeek V4 Flash (DeepInfra)
$0.09 / $0.18 per 1M tokens
DeepSeek V4 Pro (DeepInfra)
$1.30 / $2.60 per 1M tokens
Llama 4 Scout (DeepInfra)
$0.1 / $0.3 per 1M tokens
Llama 4 Maverick (DeepInfra)
$0.2 / $0.8 per 1M tokens
Qwen3 32B (DeepInfra)
$0.08 / $0.28 per 1M tokens
Mistral Nemo (DeepInfra)
$0.019 / $0.03 per 1M tokens
Gemma 4 31B Turbo (DeepInfra)
$0.09 / $0.34 per 1M tokens
About DeepInfra
Low-cost serverless inference for open models with transparent per-token rates. DeepInfra appears on CostUse as a provider of AI models with per-token pricing. This page groups the published catalog calculators so you can compare list rates, project monthly volume, and open each model with the public pricing source next to the figure.
There are 9 published models from DeepInfra here. On a 70/30 mix, combined list rates run from $0.0223 to $1.69 per 1 million tokens, with a median of $0.16. The largest listed context window in this group is 1,024,000 tokens. The most recent price check in this group is 2026-07-30.
Official DeepInfra pricing page
DeepInfra model pricing table
Figures below are list prices per 1 million tokens. The 70/30 mix is a planning default—switch the mix on each model calculator if your traffic is output-heavy.
| Model | Input / 1M | Output / 1M | Mixed 70/30 | Context | Verified |
|---|---|---|---|---|---|
| Mistral Nemo (DeepInfra) | $0.019 | $0.03 | $0.0223 | 128,000 | |
| Llama 3.1 8B Turbo (DeepInfra) | $0.02 | $0.04 | $0.026 | 128,000 | |
| DeepSeek V4 Flash (DeepInfra) | $0.09 | $0.18 | $0.117 | 1,024,000 | |
| Qwen3 32B (DeepInfra) | $0.08 | $0.28 | $0.14 | 40,000 | |
| Llama 4 Scout (DeepInfra) | $0.1 | $0.3 | $0.16 | 320,000 | |
| Gemma 4 31B Turbo (DeepInfra) | $0.09 | $0.34 | $0.165 | 256,000 | |
| Llama 3.3 70B Turbo (DeepInfra) | $0.1 | $0.32 | $0.166 | 128,000 | |
| Llama 4 Maverick (DeepInfra) | $0.2 | $0.8 | $0.38 | 1,024,000 | |
| DeepSeek V4 Pro (DeepInfra) | $1.30 | $2.60 | $1.69 | 1,024,000 |
How to read DeepInfra API cost
Input and output almost never share the same rate. On chat and agent workloads, completion length often dominates the bill. Use each model’s volume bands (from a handful of tokens to tens of millions) to find where a mini or flash tier beats the flagship on cost per accepted result—not only on unit price.
When the source publishes cached input, the model calculator can switch to that rate. For stable system prompts and RAG with repeated documents, cache often moves monthly spend more than swapping models inside the same family.
DeepInfra calculators in this catalog
The cards at the top of this page open each model’s full page: list price, interactive calculator, cost profile, and FAQs. Start with the scenario closest to your product, then compare across the rest of the DeepInfra lineup.
Frequently asked questions about DeepInfra
How many DeepInfra models have calculators on CostUse?
There are 9 published models from DeepInfra, each with input and output rates (when the source splits them), volume bands, and a link to the official pricing page.
Which DeepInfra model has the lowest list cost in this catalog?
On a 70% input / 30% output mix, Mistral Nemo (DeepInfra) is the lowest-priced option in this group at $0.0223 per 1 million tokens. Open the model calculator to stress-test your own mix.
Do DeepInfra prices include discounts and taxes?
No. Calculators use public list rates. Annual commits, promo credits, regions, taxes, and commercial discounts are not baked into the figure—always confirm the source linked on the model card.
How should I plan a monthly budget with the DeepInfra API?
Pick a model, estimate input and output tokens per request, multiply by requests per day and by 30. Compare flagship tiers with cheaper ones: classification and routing often clear quality bars on smaller models at a fraction of the cost.
When is the more expensive DeepInfra tier worth it?
DeepSeek V4 Pro (DeepInfra) sits near $1.69 per 1M tokens on a 70/30 mix, versus $0.0223 on Mistral Nemo (DeepInfra). Use the high tier for long-horizon reasoning or fewer retries; keep the cheap tier for high volume and simple tasks.