Together AI models

Open-model serverless inference, image, audio, and dedicated GPU endpoints.

https://www.together.ai/pricing

Together AI

Open-model serverless inference, image, audio, and dedicated GPU endpoints.

There are 10 published models here. On a 70/30 mix, combined list rates run from $0.078 to $6.60 per 1 million tokens, with a median of $0.805. The largest listed context window in this group is 1,000,000 tokens. The most recent price check in this group is 2026-08-14.

Official Together AI pricing page

Together AI list rates

USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.

ModelInput / 1MOutput / 1MMixed 70/30ContextVerified
Gemma 3n E4B Instruct (Together) $0.06 $0.12 $0.078 32,000
GPT OSS 20B (Together) $0.05 $0.2 $0.095 131,072
Qwen3.5 9B (Together) $0.17 $0.25 $0.194 128,000
GPT OSS 120B (Together) $0.15 $0.6 $0.285 131,072
MiniMax M3 (Together) $0.3 $1.20 $0.57 1,000,000
Llama 3.3 70B (Together) $1.04 $1.04 $1.04 131,072
Kimi K2.6 (Together) $1.20 $4.50 $2.19 256,000
DeepSeek V4 Pro (Together) $1.74 $3.48 $2.262 512,000
GLM 5.2 (Together) $1.40 $4.40 $2.30 200,000
Kimi K3 (Together) $3.00 $15.00 $6.60 256,000

Questions about Together AI rates

Do these Together AI figures include discounts and tax?

No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.

When is the more expensive Together AI tier worth the extra?

Kimi K3 (Together) sits near $6.60 per 1M tokens on a 70/30 mix, versus $0.078 on Gemma 3n E4B Instruct (Together). Use the high tier when retries or long reasoning actually fail on the cheap one.