DeepInfra

Gemma 4 31B Turbo (DeepInfra) token cost

Gemma 4 31B Turbo on DeepInfra — compare speed hosts like Cerebras at different unit rates.

Published list rates

Input

$0.09

per 1M tokens

Output

$0.34

per 1M tokens

Context window

256,000

tokens

Pricing source: DeepInfra pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Gemma 4 31B Turbo (DeepInfra) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000009 $0.00000034 $0.000000165
10 $0.0000009 $0.000003 $0.000002
100 $0.000009 $0.000034 $0.000017
200 $0.000018 $0.000068 $0.000033
300 $0.000027 $0.000102 $0.00005
400 $0.000036 $0.000136 $0.000066
500 $0.000045 $0.00017 $0.000083
1,000 $0.00009 $0.00034 $0.000165
2,000 $0.00018 $0.00068 $0.00033
5,000 $0.00045 $0.0017 $0.000825
10,000 $0.0009 $0.0034 $0.00165
25,000 $0.00225 $0.0085 $0.004125
50,000 $0.0045 $0.017 $0.00825
100,000 $0.009 $0.034 $0.0165
250,000 $0.0225 $0.085 $0.04125
500,000 $0.045 $0.17 $0.0825
1,000,000 $0.09 $0.34 $0.165
5,000,000 $0.45 $1.70 $0.825
10,000,000 $0.9 $3.40 $1.65
50,000,000 $4.50 $17.00 $8.25
100,000,000 $9.00 $34.00 $16.50
500,000,000 $45.00 $170.00 $82.50
1,000,000,000 $90.00 $340.00 $165.00
10,000,000,000 $900.00 $3,400.00 $1,650.00
100,000,000,000 $9,000.00 $34,000.00 $16,500.00

Tokens for a fixed budget

If you cap spend, this is roughly how many Gemma 4 31B Turbo (DeepInfra) tokens that money buys.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 111,111 29,412 60,606
$0.1 1.11M 294,118 606,061
$0.5 5.56M 1.47M 3.03M
$1.00 11.1M 2.94M 6.06M
$5.00 55.6M 14.7M 30.3M
$10.00 111.1M 29.4M 60.6M
$20.00 222.2M 58.8M 121.2M
$25.00 277.8M 73.5M 151.5M
$30.00 333.3M 88.2M 181.8M
$50.00 555.6M 147.1M 303M
$100.00 1.11B 294.1M 606.1M
$250.00 2.78B 735.3M 1.52B
$500.00 5.56B 1.47B 3.03B
$1,000.00 11.1B 2.94B 6.06B
$5,000.00 55.6B 14.7B 30.3B
$10,000.00 111.1B 29.4B 60.6B

At these rates

  • A workload with 70% input and 30% output costs $0.165 per 1 million total tokens and $16.50 per 100 million.
  • At the same token count, output costs 3.78 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 256,000-token context window with input alone would cost about $0.02304, before any output.

DeepInfra

Low-cost serverless inference for open models with transparent per-token rates.

All DeepInfra models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models