DeepInfra

DeepSeek V4 Flash (DeepInfra) token cost

DeepSeek V4 Flash on DeepInfra with cache discounts — compare against first-party and Fireworks.

Published list rates

Input

$0.09

per 1M tokens

Output

$0.18

per 1M tokens

Cached input

$0.018

per 1M tokens

Context window

1,024,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

What DeepSeek V4 Flash (DeepInfra) costs if you only send tokens, only get a reply, or do both (70% in / 30% out). Rows start small and go up to huge traffic.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000009 $0.00000018 $0.000000117
10 $0.0000009 $0.000002 $0.000001
100 $0.000009 $0.000018 $0.000012
200 $0.000018 $0.000036 $0.000023
300 $0.000027 $0.000054 $0.000035
400 $0.000036 $0.000072 $0.000047
500 $0.000045 $0.00009 $0.000059
1,000 $0.00009 $0.00018 $0.000117
2,000 $0.00018 $0.00036 $0.000234
5,000 $0.00045 $0.0009 $0.000585
10,000 $0.0009 $0.0018 $0.00117
25,000 $0.00225 $0.0045 $0.002925
50,000 $0.0045 $0.009 $0.00585
100,000 $0.009 $0.018 $0.0117
250,000 $0.0225 $0.045 $0.02925
500,000 $0.045 $0.09 $0.0585
1,000,000 $0.09 $0.18 $0.117
5,000,000 $0.45 $0.9 $0.585
10,000,000 $0.9 $1.80 $1.17
50,000,000 $4.50 $9.00 $5.85
100,000,000 $9.00 $18.00 $11.70
500,000,000 $45.00 $90.00 $58.50
1,000,000,000 $90.00 $180.00 $117.00
10,000,000,000 $900.00 $1,800.00 $1,170.00
100,000,000,000 $9,000.00 $18,000.00 $11,700.00

Tokens for a fixed budget

Tokens DeepSeek V4 Flash (DeepInfra) can process for a set budget — send-only, reply-only, and a 70/30 mix.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 111,111 55,556 85,470
$0.1 1.11M 555,556 854,701
$0.5 5.56M 2.78M 4.27M
$1.00 11.1M 5.56M 8.55M
$5.00 55.6M 27.8M 42.7M
$10.00 111.1M 55.6M 85.5M
$20.00 222.2M 111.1M 170.9M
$25.00 277.8M 138.9M 213.7M
$30.00 333.3M 166.7M 256.4M
$50.00 555.6M 277.8M 427.4M
$100.00 1.11B 555.6M 854.7M
$250.00 2.78B 1.39B 2.14B
$500.00 5.56B 2.78B 4.27B
$1,000.00 11.1B 5.56B 8.55B
$5,000.00 55.6B 27.8B 42.7B
$10,000.00 111.1B 55.6B 85.5B

At these rates

  • A workload with 70% input and 30% output costs $0.117 per 1 million total tokens and $11.70 per 100 million.
  • At the same token count, output costs 2 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 1,024,000-token context window with input alone would cost about $0.09216, before any output.
  • Cached input is $0.018 per million, 80% below standard input when the workload meets the provider’s cache rules.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models