DeepInfra

Llama 3.1 8B Turbo (DeepInfra) token cost

Ultra-cheap Llama 3.1 8B Turbo on DeepInfra for routing, classification, and light agents.

Published list rates

Input

$0.02

per 1M tokens

Output

$0.04

per 1M tokens

Context window

128,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Llama 3.1 8B Turbo (DeepInfra) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000002 $0.00000004 $0.000000026
10 $0.0000002 $0.0000004 $0.00000026
100 $0.000002 $0.000004 $0.000003
200 $0.000004 $0.000008 $0.000005
300 $0.000006 $0.000012 $0.000008
400 $0.000008 $0.000016 $0.00001
500 $0.00001 $0.00002 $0.000013
1,000 $0.00002 $0.00004 $0.000026
2,000 $0.00004 $0.00008 $0.000052
5,000 $0.0001 $0.0002 $0.00013
10,000 $0.0002 $0.0004 $0.00026
25,000 $0.0005 $0.001 $0.00065
50,000 $0.001 $0.002 $0.0013
100,000 $0.002 $0.004 $0.0026
250,000 $0.005 $0.01 $0.0065
500,000 $0.01 $0.02 $0.013
1,000,000 $0.02 $0.04 $0.026
5,000,000 $0.1 $0.2 $0.13
10,000,000 $0.2 $0.4 $0.26
50,000,000 $1.00 $2.00 $1.30
100,000,000 $2.00 $4.00 $2.60
500,000,000 $10.00 $20.00 $13.00
1,000,000,000 $20.00 $40.00 $26.00
10,000,000,000 $200.00 $400.00 $260.00
100,000,000,000 $2,000.00 $4,000.00 $2,600.00

Tokens for a fixed budget

If you cap spend, this is roughly how many Llama 3.1 8B Turbo (DeepInfra) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 500,000 250,000 384,615
$0.1 5M 2.5M 3.85M
$0.5 25M 12.5M 19.2M
$1.00 50M 25M 38.5M
$5.00 250M 125M 192.3M
$10.00 500M 250M 384.6M
$20.00 1B 500M 769.2M
$25.00 1.25B 625M 961.5M
$30.00 1.5B 750M 1.15B
$50.00 2.5B 1.25B 1.92B
$100.00 5B 2.5B 3.85B
$250.00 12.5B 6.25B 9.62B
$500.00 25B 12.5B 19.2B
$1,000.00 50B 25B 38.5B
$5,000.00 250B 125B 192.3B
$10,000.00 500B 250B 384.6B

At these rates

  • A workload with 70% input and 30% output costs $0.026 per 1 million total tokens and $2.60 per 100 million.
  • At the same token count, output costs 2 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 128,000-token context window with input alone would cost about $0.00256, before any output.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models