DeepInfra

Llama 4 Scout (DeepInfra) token cost

Meta Llama 4 Scout multimodal tier on DeepInfra for efficient long-context work.

Published list rates

Input

$0.1

per 1M tokens

Output

$0.3

per 1M tokens

Context window

320,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Llama 4 Scout (DeepInfra) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.0000001 $0.0000003 $0.00000016
10 $0.000001 $0.000003 $0.000002
100 $0.00001 $0.00003 $0.000016
200 $0.00002 $0.00006 $0.000032
300 $0.00003 $0.00009 $0.000048
400 $0.00004 $0.00012 $0.000064
500 $0.00005 $0.00015 $0.00008
1,000 $0.0001 $0.0003 $0.00016
2,000 $0.0002 $0.0006 $0.00032
5,000 $0.0005 $0.0015 $0.0008
10,000 $0.001 $0.003 $0.0016
25,000 $0.0025 $0.0075 $0.004
50,000 $0.005 $0.015 $0.008
100,000 $0.01 $0.03 $0.016
250,000 $0.025 $0.075 $0.04
500,000 $0.05 $0.15 $0.08
1,000,000 $0.1 $0.3 $0.16
5,000,000 $0.5 $1.50 $0.8
10,000,000 $1.00 $3.00 $1.60
50,000,000 $5.00 $15.00 $8.00
100,000,000 $10.00 $30.00 $16.00
500,000,000 $50.00 $150.00 $80.00
1,000,000,000 $100.00 $300.00 $160.00
10,000,000,000 $1,000.00 $3,000.00 $1,600.00
100,000,000,000 $10,000.00 $30,000.00 $16,000.00

Tokens for a fixed budget

If you cap spend, this is roughly how many Llama 4 Scout (DeepInfra) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 100,000 33,333 62,500
$0.1 1M 333,333 625,000
$0.5 5M 1.67M 3.13M
$1.00 10M 3.33M 6.25M
$5.00 50M 16.7M 31.3M
$10.00 100M 33.3M 62.5M
$20.00 200M 66.7M 125M
$25.00 250M 83.3M 156.3M
$30.00 300M 100M 187.5M
$50.00 500M 166.7M 312.5M
$100.00 1B 333.3M 625M
$250.00 2.5B 833.3M 1.56B
$500.00 5B 1.67B 3.13B
$1,000.00 10B 3.33B 6.25B
$5,000.00 50B 16.7B 31.3B
$10,000.00 100B 33.3B 62.5B

At these rates

  • A workload with 70% input and 30% output costs $0.16 per 1 million total tokens and $16.00 per 100 million.
  • At the same token count, output costs 3 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 320,000-token context window with input alone would cost about $0.032, before any output.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models