DeepInfra
Gemma 4 31B Turbo (DeepInfra) token cost
Gemma 4 31B Turbo on DeepInfra — compare speed hosts like Cerebras at different unit rates.

Published list rates
Input
$0.09
per 1M tokens
Output
$0.34
per 1M tokens
Context window
256,000
tokens
Pricing source: DeepInfra pricing
Estimate cost
Estimate from text
Calculate from a token total
Project monthly volume
Daily requests × tokens per request × 30 days.
Estimated cost
- Estimated tokens
- —
- Characters
- —
- Words
- —
- Input tokens
- —
- Output tokens
- —
- Input cost
- —
- Output cost
- —
- Total cost
- —
- Combined price per 1M tokens
- —
- Estimated monthly cost
- —
- Estimated daily cost
- —
Cost as volume grows
Gemma 4 31B Turbo (DeepInfra) at rising volume: prompt-only, reply-only, and a mixed chat (70% in, 30% out).
| Tokens | Input only | Output only | Mix (70% in / 30% out) |
|---|---|---|---|
| 1 | $0.00000009 | $0.00000034 | $0.000000165 |
| 10 | $0.0000009 | $0.000003 | $0.000002 |
| 100 | $0.000009 | $0.000034 | $0.000017 |
| 200 | $0.000018 | $0.000068 | $0.000033 |
| 300 | $0.000027 | $0.000102 | $0.00005 |
| 400 | $0.000036 | $0.000136 | $0.000066 |
| 500 | $0.000045 | $0.00017 | $0.000083 |
| 1,000 | $0.00009 | $0.00034 | $0.000165 |
| 2,000 | $0.00018 | $0.00068 | $0.00033 |
| 5,000 | $0.00045 | $0.0017 | $0.000825 |
| 10,000 | $0.0009 | $0.0034 | $0.00165 |
| 25,000 | $0.00225 | $0.0085 | $0.004125 |
| 50,000 | $0.0045 | $0.017 | $0.00825 |
| 100,000 | $0.009 | $0.034 | $0.0165 |
| 250,000 | $0.0225 | $0.085 | $0.04125 |
| 500,000 | $0.045 | $0.17 | $0.0825 |
| 1,000,000 | $0.09 | $0.34 | $0.165 |
| 5,000,000 | $0.45 | $1.70 | $0.825 |
| 10,000,000 | $0.9 | $3.40 | $1.65 |
| 50,000,000 | $4.50 | $17.00 | $8.25 |
| 100,000,000 | $9.00 | $34.00 | $16.50 |
| 500,000,000 | $45.00 | $170.00 | $82.50 |
| 1,000,000,000 | $90.00 | $340.00 | $165.00 |
| 10,000,000,000 | $900.00 | $3,400.00 | $1,650.00 |
| 100,000,000,000 | $9,000.00 | $34,000.00 | $16,500.00 |
Tokens for a fixed budget
If you cap spend, this is roughly how many Gemma 4 31B Turbo (DeepInfra) tokens that money buys.
| Budget | Input tokens | Output tokens | Mixed tokens (70/30) |
|---|---|---|---|
| $0.01 | 111,111 | 29,412 | 60,606 |
| $0.1 | 1.11M | 294,118 | 606,061 |
| $0.5 | 5.56M | 1.47M | 3.03M |
| $1.00 | 11.1M | 2.94M | 6.06M |
| $5.00 | 55.6M | 14.7M | 30.3M |
| $10.00 | 111.1M | 29.4M | 60.6M |
| $20.00 | 222.2M | 58.8M | 121.2M |
| $25.00 | 277.8M | 73.5M | 151.5M |
| $30.00 | 333.3M | 88.2M | 181.8M |
| $50.00 | 555.6M | 147.1M | 303M |
| $100.00 | 1.11B | 294.1M | 606.1M |
| $250.00 | 2.78B | 735.3M | 1.52B |
| $500.00 | 5.56B | 1.47B | 3.03B |
| $1,000.00 | 11.1B | 2.94B | 6.06B |
| $5,000.00 | 55.6B | 14.7B | 30.3B |
| $10,000.00 | 111.1B | 29.4B | 60.6B |
At these rates
- A workload with 70% input and 30% output costs $0.165 per 1 million total tokens and $16.50 per 100 million.
- At the same token count, output costs 3.78 times the input rate. Response length therefore deserves its own budget limit.
- Filling the 256,000-token context window with input alone would cost about $0.02304, before any output.
These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.