Groq models

Ultra-fast inference for open models with transparent per-token pricing.

https://groq.com/pricing

Groq

Ultra-fast inference for open models with transparent per-token pricing.

There are 7 published models here. On a 70/30 mix, combined list rates run from $0.03 to $1.32 per 1 million tokens, with a median of $0.1425. The largest listed context window in this group is 131,072 tokens. The most recent price check in this group is 2026-08-19.

Official Groq pricing page

Groq list rates

USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.

ModelInput / 1MOutput / 1MMixed 70/30ContextVerified
Llama Prompt Guard 2 22M (Groq) $0.03 $0.03 $0.03 512
Llama Prompt Guard 2 86M (Groq) $0.04 $0.04 $0.04 512
Llama 3.1 8B Instant (Groq) $0.05 $0.08 $0.059 131,072
GPT OSS 20B (Groq) $0.075 $0.3 $0.1425 131,072
GPT OSS 120B (Groq) $0.15 $0.6 $0.285 131,072
Llama 3.3 70B Versatile (Groq) $0.59 $0.79 $0.65 131,072
Qwen 3.6 27B (Groq) $0.6 $3.00 $1.32 131,072

Questions about Groq rates

Do these Groq figures include discounts and tax?

No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.

When is the more expensive Groq tier worth the extra?

Qwen 3.6 27B (Groq) sits near $1.32 per 1M tokens on a 70/30 mix, versus $0.03 on Llama Prompt Guard 2 22M (Groq). Use the high tier when retries or long reasoning actually fail on the cheap one.