DeepInfra

Llama 4 Maverick (DeepInfra) token cost

Llama 4 Maverick on DeepInfra for higher-quality multimodal generation at open-host rates.

Published list rates

Input

$0.2

per 1M tokens

Output

$0.8

per 1M tokens

Context window

1,024,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Three views of Llama 4 Maverick (DeepInfra): sending only, receiving only, and a 70/30 mix. Same list rates, bigger rows.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.0000002 $0.0000008 $0.00000038
10 $0.000002 $0.000008 $0.000004
100 $0.00002 $0.00008 $0.000038
200 $0.00004 $0.00016 $0.000076
300 $0.00006 $0.00024 $0.000114
400 $0.00008 $0.00032 $0.000152
500 $0.0001 $0.0004 $0.00019
1,000 $0.0002 $0.0008 $0.00038
2,000 $0.0004 $0.0016 $0.00076
5,000 $0.001 $0.004 $0.0019
10,000 $0.002 $0.008 $0.0038
25,000 $0.005 $0.02 $0.0095
50,000 $0.01 $0.04 $0.019
100,000 $0.02 $0.08 $0.038
250,000 $0.05 $0.2 $0.095
500,000 $0.1 $0.4 $0.19
1,000,000 $0.2 $0.8 $0.38
5,000,000 $1.00 $4.00 $1.90
10,000,000 $2.00 $8.00 $3.80
50,000,000 $10.00 $40.00 $19.00
100,000,000 $20.00 $80.00 $38.00
500,000,000 $100.00 $400.00 $190.00
1,000,000,000 $200.00 $800.00 $380.00
10,000,000,000 $2,000.00 $8,000.00 $3,800.00
100,000,000,000 $20,000.00 $80,000.00 $38,000.00

Tokens for a fixed budget

Tokens Llama 4 Maverick (DeepInfra) can process for a set budget — send-only, reply-only, and a 70/30 mix.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 50,000 12,500 26,316
$0.1 500,000 125,000 263,158
$0.5 2.5M 625,000 1.32M
$1.00 5M 1.25M 2.63M
$5.00 25M 6.25M 13.2M
$10.00 50M 12.5M 26.3M
$20.00 100M 25M 52.6M
$25.00 125M 31.3M 65.8M
$30.00 150M 37.5M 78.9M
$50.00 250M 62.5M 131.6M
$100.00 500M 125M 263.2M
$250.00 1.25B 312.5M 657.9M
$500.00 2.5B 625M 1.32B
$1,000.00 5B 1.25B 2.63B
$5,000.00 25B 6.25B 13.2B
$10,000.00 50B 12.5B 26.3B

At these rates

  • A workload with 70% input and 30% output costs $0.38 per 1 million total tokens and $38.00 per 100 million.
  • At the same token count, output costs 4 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 1,024,000-token context window with input alone would cost about $0.2048, before any output.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models