
Llama 3.3 70B (Hyperbolic)
$0.12 in · $0.3 out per 1M tokens
Pick a provider, then a model. Rates, source, and check date are on the model page.
Open-model inference API plus GPU marketplace with aggressive per-token rates.

$0.12 in · $0.3 out per 1M tokens

$0.1 in · $0.1 out per 1M tokens

$0.25 in · $0.25 out per 1M tokens
SambaNova Cloud inference with high-throughput open models and public per-token rates.

$0.22 in · $0.59 out per 1M tokens

$0.38 in · $1.15 out per 1M tokens

$0.6 in · $1.20 out per 1M tokens

$0.6 in · $2.40 out per 1M tokens
Alibaba Cloud Model Studio first-party Qwen models with international token rates.

$2.50 in · $7.50 out per 1M tokens

$2.00 in · $6.00 out per 1M tokens

$0.4 in · $2.40 out per 1M tokens

$0.1 in · $0.4 out per 1M tokens
Fast open-model Model APIs with cache-aware per-token pricing.

$1.40 in · $4.40 out per 1M tokens

$0.5 in · $1.50 out per 1M tokens

$0.14 in · $0.4 out per 1M tokens

$0.3 in · $1.20 out per 1M tokens
Muse Spark multimodal reasoning models on the Meta Model API, billed per token.

$1.25 in · $4.25 out per 1M tokens

$0.1 in · $0.2 out per 1M tokens
First-party Microsoft AI (MAI) reasoning and coding models on Microsoft Foundry.

$2.00 in · $8.00 out per 1M tokens

$0.2 in · $1.20 out per 1M tokens

$0.75 in · $4.50 out per 1M tokens
Ling open-weight hybrid reasoning models served through third-party serverless APIs.

$0.06 in · $0.18 out per 1M tokens
Spark X models from iFLYTEK's open platform, served on domestic Chinese compute.

$0.236 in · $0.886 out per 1M tokens
Nex-N2 open agentic models from the Shanghai Innovation Institute alliance, via API hosts.

$0.025 in · $0.1 out per 1M tokens

$0.25 in · $1.00 out per 1M tokens
Unified multi-provider model router with public per-token list rates.

$0.75 in · $4.50 out per 1M tokens

$3.00 in · $15.00 out per 1M tokens

$0.75 in · $3.75 out per 1M tokens

$0.15 in · $0.6 out per 1M tokens

$0.269 in · $0.4 out per 1M tokens

$0.1 in · $0.3 out per 1M tokens