InclusionAI

Ling 3.0 Flash token cost

Ling 3.0 Flash: 124B MoE with 5.1B active per token, at $0.06/$0.18 per 1M on Novita serverless.

Published list rates

Input

$0.06

per 1M tokens

Output

$0.18

per 1M tokens

Cached input

$0.012

per 1M tokens

Context window

256,000

tokens

Pricing source: Novita serverless pricing

Estimate cost

Estimate from text

Calculate from a token total

Project monthly volume

Daily requests × tokens per request × 30 days.

Estimated cost

Estimated tokens
Characters
Words
Input tokens
Output tokens
Input cost
Output cost
Total cost
Combined price per 1M tokens
Estimated monthly cost
Estimated daily cost

Cost as volume grows

Read down to see Ling 3.0 Flash get more expensive as tokens grow. The last column is a typical mix: 70% input, 30% output.

Tokens Input only Output only Mix (70% in / 30% out)
1 $0.00000006 $0.00000018 $0.000000096
10 $0.0000006 $0.000002 $0.00000096
100 $0.000006 $0.000018 $0.00001
200 $0.000012 $0.000036 $0.000019
300 $0.000018 $0.000054 $0.000029
400 $0.000024 $0.000072 $0.000038
500 $0.00003 $0.00009 $0.000048
1,000 $0.00006 $0.00018 $0.000096
2,000 $0.00012 $0.00036 $0.000192
5,000 $0.0003 $0.0009 $0.00048
10,000 $0.0006 $0.0018 $0.00096
25,000 $0.0015 $0.0045 $0.0024
50,000 $0.003 $0.009 $0.0048
100,000 $0.006 $0.018 $0.0096
250,000 $0.015 $0.045 $0.024
500,000 $0.03 $0.09 $0.048
1,000,000 $0.06 $0.18 $0.096
5,000,000 $0.3 $0.9 $0.48
10,000,000 $0.6 $1.80 $0.96
50,000,000 $3.00 $9.00 $4.80
100,000,000 $6.00 $18.00 $9.60
500,000,000 $30.00 $90.00 $48.00
1,000,000,000 $60.00 $180.00 $96.00
10,000,000,000 $600.00 $1,800.00 $960.00
100,000,000,000 $6,000.00 $18,000.00 $9,600.00

Tokens for a fixed budget

Tokens Ling 3.0 Flash can process for a set budget — send-only, reply-only, and a 70/30 mix.

Budget Input tokens Output tokens Mixed tokens (70/30)
$0.01 166,667 55,556 104,167
$0.1 1.67M 555,556 1.04M
$0.5 8.33M 2.78M 5.21M
$1.00 16.7M 5.56M 10.4M
$5.00 83.3M 27.8M 52.1M
$10.00 166.7M 55.6M 104.2M
$20.00 333.3M 111.1M 208.3M
$25.00 416.7M 138.9M 260.4M
$30.00 500M 166.7M 312.5M
$50.00 833.3M 277.8M 520.8M
$100.00 1.67B 555.6M 1.04B
$250.00 4.17B 1.39B 2.6B
$500.00 8.33B 2.78B 5.21B
$1,000.00 16.7B 5.56B 10.4B
$5,000.00 83.3B 27.8B 52.1B
$10,000.00 166.7B 55.6B 104.2B

At these rates

  • A workload with 70% input and 30% output costs $0.096 per 1 million total tokens and $9.60 per 100 million.
  • At the same token count, output costs 3 times the input rate. Response length therefore deserves its own budget limit.
  • Filling the 256,000-token context window with input alone would cost about $0.01536, before any output.
  • Cached input is $0.012 per million, 80% below standard input when the workload meets the provider’s cache rules.

InclusionAI

Ling open-weight hybrid reasoning models served through third-party serverless APIs.

All InclusionAI models

MIT-licensed open weights. Rates shown are Novita's Standard serverless list; OpenRouter lists $0.021/$0.063 per 1M. Cache reads are $0.012 per 1M on Novita.

These figures use public list prices. Cache eligibility, batch or priority modes, long-context surcharges, discounts, and regional billing can change the invoice.

Related models