DeepInfra

Qwen3 32B (DeepInfra) token cost

Qwen3 32B on DeepInfra for multilingual chat and coding at low unit cost.

Published list rates

Input

$0.08

per 1M tokens

Output

$0.28

per 1M tokens

Context window

40,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Three views of Qwen3 32B (DeepInfra): sending only, receiving only, and a 70/30 mix. Same list rates, bigger rows.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000008 $0.00000028 $0.00000014
10 $0.0000008 $0.000003 $0.000001
100 $0.000008 $0.000028 $0.000014
200 $0.000016 $0.000056 $0.000028
300 $0.000024 $0.000084 $0.000042
400 $0.000032 $0.000112 $0.000056
500 $0.00004 $0.00014 $0.00007
1,000 $0.00008 $0.00028 $0.00014
2,000 $0.00016 $0.00056 $0.00028
5,000 $0.0004 $0.0014 $0.0007
10,000 $0.0008 $0.0028 $0.0014
25,000 $0.002 $0.007 $0.0035
50,000 $0.004 $0.014 $0.007
100,000 $0.008 $0.028 $0.014
250,000 $0.02 $0.07 $0.035
500,000 $0.04 $0.14 $0.07
1,000,000 $0.08 $0.28 $0.14
5,000,000 $0.4 $1.40 $0.7
10,000,000 $0.8 $2.80 $1.40
50,000,000 $4.00 $14.00 $7.00
100,000,000 $8.00 $28.00 $14.00
500,000,000 $40.00 $140.00 $70.00
1,000,000,000 $80.00 $280.00 $140.00
10,000,000,000 $800.00 $2,800.00 $1,400.00
100,000,000,000 $8,000.00 $28,000.00 $14,000.00

Tokens for a fixed budget

How far a fixed spend goes on Qwen3 32B (DeepInfra) at these list rates.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 125,000 35,714 71,429
$0.1 1.25M 357,143 714,286
$0.5 6.25M 1.79M 3.57M
$1.00 12.5M 3.57M 7.14M
$5.00 62.5M 17.9M 35.7M
$10.00 125M 35.7M 71.4M
$20.00 250M 71.4M 142.9M
$25.00 312.5M 89.3M 178.6M
$30.00 375M 107.1M 214.3M
$50.00 625M 178.6M 357.1M
$100.00 1.25B 357.1M 714.3M
$250.00 3.13B 892.9M 1.79B
$500.00 6.25B 1.79B 3.57B
$1,000.00 12.5B 3.57B 7.14B
$5,000.00 62.5B 17.9B 35.7B
$10,000.00 125B 35.7B 71.4B

At these rates

  • A workload with 70% input and 30% output costs $0.14 per 1 million total tokens and $14.00 per 100 million.
  • At the same token count, output costs 3.5 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 40,000-token context window with input alone would cost about $0.0032, before any output.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models