DeepInfra

DeepSeek V4 Pro (DeepInfra) token cost

DeepSeek V4 Pro hosted on DeepInfra for long-context reasoning with cache-aware rates.

Published list rates

Input

$1.30

per 1M tokens

Output

$2.60

per 1M tokens

Cached input

$0.1

per 1M tokens

Context window

1,024,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

DeepSeek V4 Pro (DeepInfra) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.000001 $0.000003 $0.000002
10 $0.000013 $0.000026 $0.000017
100 $0.00013 $0.00026 $0.000169
200 $0.00026 $0.00052 $0.000338
300 $0.00039 $0.00078 $0.000507
400 $0.00052 $0.00104 $0.000676
500 $0.00065 $0.0013 $0.000845
1,000 $0.0013 $0.0026 $0.00169
2,000 $0.0026 $0.0052 $0.00338
5,000 $0.0065 $0.013 $0.00845
10,000 $0.013 $0.026 $0.0169
25,000 $0.0325 $0.065 $0.04225
50,000 $0.065 $0.13 $0.0845
100,000 $0.13 $0.26 $0.169
250,000 $0.325 $0.65 $0.4225
500,000 $0.65 $1.30 $0.845
1,000,000 $1.30 $2.60 $1.69
5,000,000 $6.50 $13.00 $8.45
10,000,000 $13.00 $26.00 $16.90
50,000,000 $65.00 $130.00 $84.50
100,000,000 $130.00 $260.00 $169.00
500,000,000 $650.00 $1,300.00 $845.00
1,000,000,000 $1,300.00 $2,600.00 $1,690.00
10,000,000,000 $13,000.00 $26,000.00 $16,900.00
100,000,000,000 $130,000.00 $260,000.00 $169,000.00

Tokens for a fixed budget

If you cap spend, this is roughly how many DeepSeek V4 Pro (DeepInfra) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 7,692 3,846 5,917
$0.1 76,923 38,462 59,172
$0.5 384,615 192,308 295,858
$1.00 769,231 384,615 591,716
$5.00 3.85M 1.92M 2.96M
$10.00 7.69M 3.85M 5.92M
$20.00 15.4M 7.69M 11.8M
$25.00 19.2M 9.62M 14.8M
$30.00 23.1M 11.5M 17.8M
$50.00 38.5M 19.2M 29.6M
$100.00 76.9M 38.5M 59.2M
$250.00 192.3M 96.2M 147.9M
$500.00 384.6M 192.3M 295.9M
$1,000.00 769.2M 384.6M 591.7M
$5,000.00 3.85B 1.92B 2.96B
$10,000.00 7.69B 3.85B 5.92B

At these rates

  • A workload with 70% input and 30% output costs $1.69 per 1 million total tokens and $169.00 per 100 million.
  • At the same token count, output costs 2 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 1,024,000-token context window with input alone would cost about $1.3312, before any output.
  • Cached input is $0.1 per million, 92% below standard input when the workload meets the provider’s cache rules.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models