DeepInfra

Llama 3.3 70B Turbo (DeepInfra) token cost

Llama 3.3 70B Turbo on DeepInfra — one of the cheapest 70B hosts for high-volume chat.

Published list rates

Input

$0.1

per 1M tokens

Output

$0.32

per 1M tokens

Context window

128,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Read down to see Llama 3.3 70B Turbo (DeepInfra) get more expensive as tokens grow. The last column is a typical mix: 70% input, 30% output.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.0000001 $0.00000032 $0.000000166
10 $0.000001 $0.000003 $0.000002
100 $0.00001 $0.000032 $0.000017
200 $0.00002 $0.000064 $0.000033
300 $0.00003 $0.000096 $0.00005
400 $0.00004 $0.000128 $0.000066
500 $0.00005 $0.00016 $0.000083
1,000 $0.0001 $0.00032 $0.000166
2,000 $0.0002 $0.00064 $0.000332
5,000 $0.0005 $0.0016 $0.00083
10,000 $0.001 $0.0032 $0.00166
25,000 $0.0025 $0.008 $0.00415
50,000 $0.005 $0.016 $0.0083
100,000 $0.01 $0.032 $0.0166
250,000 $0.025 $0.08 $0.0415
500,000 $0.05 $0.16 $0.083
1,000,000 $0.1 $0.32 $0.166
5,000,000 $0.5 $1.60 $0.83
10,000,000 $1.00 $3.20 $1.66
50,000,000 $5.00 $16.00 $8.30
100,000,000 $10.00 $32.00 $16.60
500,000,000 $50.00 $160.00 $83.00
1,000,000,000 $100.00 $320.00 $166.00
10,000,000,000 $1,000.00 $3,200.00 $1,660.00
100,000,000,000 $10,000.00 $32,000.00 $16,600.00

Tokens for a fixed budget

Tokens Llama 3.3 70B Turbo (DeepInfra) can process for a set budget — send-only, reply-only, and a 70/30 mix.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 100,000 31,250 60,241
$0.1 1M 312,500 602,410
$0.5 5M 1.56M 3.01M
$1.00 10M 3.13M 6.02M
$5.00 50M 15.6M 30.1M
$10.00 100M 31.3M 60.2M
$20.00 200M 62.5M 120.5M
$25.00 250M 78.1M 150.6M
$30.00 300M 93.8M 180.7M
$50.00 500M 156.3M 301.2M
$100.00 1B 312.5M 602.4M
$250.00 2.5B 781.3M 1.51B
$500.00 5B 1.56B 3.01B
$1,000.00 10B 3.13B 6.02B
$5,000.00 50B 15.6B 30.1B
$10,000.00 100B 31.3B 60.2B

At these rates

  • A workload with 70% input and 30% output costs $0.166 per 1 million total tokens and $16.60 per 100 million.
  • At the same token count, output costs 3.2 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 128,000-token context window with input alone would cost about $0.0128, before any output.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models