Baseten

GPT OSS 120B (Baseten) token cost

OpenAI gpt-oss-120b on Baseten Model APIs at low open-weight token rates.

Published list rates

Input

$0.1

per 1M tokens

Output

$0.5

per 1M tokens

Context window

128,000

tokens

Pricing source: Baseten Model APIs pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Three views of GPT OSS 120B (Baseten): sending only, receiving only, and a 70/30 mix. Same list rates, bigger rows.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.0000001 $0.0000005 $0.00000022
10 $0.000001 $0.000005 $0.000002
100 $0.00001 $0.00005 $0.000022
200 $0.00002 $0.0001 $0.000044
300 $0.00003 $0.00015 $0.000066
400 $0.00004 $0.0002 $0.000088
500 $0.00005 $0.00025 $0.00011
1,000 $0.0001 $0.0005 $0.00022
2,000 $0.0002 $0.001 $0.00044
5,000 $0.0005 $0.0025 $0.0011
10,000 $0.001 $0.005 $0.0022
25,000 $0.0025 $0.0125 $0.0055
50,000 $0.005 $0.025 $0.011
100,000 $0.01 $0.05 $0.022
250,000 $0.025 $0.125 $0.055
500,000 $0.05 $0.25 $0.11
1,000,000 $0.1 $0.5 $0.22
5,000,000 $0.5 $2.50 $1.10
10,000,000 $1.00 $5.00 $2.20
50,000,000 $5.00 $25.00 $11.00
100,000,000 $10.00 $50.00 $22.00
500,000,000 $50.00 $250.00 $110.00
1,000,000,000 $100.00 $500.00 $220.00
10,000,000,000 $1,000.00 $5,000.00 $2,200.00
100,000,000,000 $10,000.00 $50,000.00 $22,000.00

Tokens for a fixed budget

Tokens GPT OSS 120B (Baseten) can process for a set budget — send-only, reply-only, and a 70/30 mix.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 100,000 20,000 45,455
$0.1 1M 200,000 454,545
$0.5 5M 1M 2.27M
$1.00 10M 2M 4.55M
$5.00 50M 10M 22.7M
$10.00 100M 20M 45.5M
$20.00 200M 40M 90.9M
$25.00 250M 50M 113.6M
$30.00 300M 60M 136.4M
$50.00 500M 100M 227.3M
$100.00 1B 200M 454.5M
$250.00 2.5B 500M 1.14B
$500.00 5B 1B 2.27B
$1,000.00 10B 2B 4.55B
$5,000.00 50B 10B 22.7B
$10,000.00 100B 20B 45.5B

At these rates

  • A workload with 70% input and 30% output costs $0.22 per 1 million total tokens and $22.00 per 100 million.
  • At the same token count, output costs 5 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 128,000-token context window with input alone would cost about $0.0128, before any output.

Baseten

Model APIs and dedicated GPU inference with public per-token list rates.

All Baseten models

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models