Groq models
Ultra-fast inference for open models with transparent per-token pricing.


Llama 3.3 70B Versatile (Groq)
$0.59 / $0.79 per 1M tokens

Llama 3.1 8B Instant (Groq)
$0.05 / $0.08 per 1M tokens

GPT OSS 120B (Groq)
$0.15 / $0.6 per 1M tokens

GPT OSS 20B (Groq)
$0.075 / $0.3 per 1M tokens

Qwen 3.6 27B (Groq)
$0.6 / $3.00 per 1M tokens

Llama Prompt Guard 2 22M (Groq)
$0.03 / $0.03 per 1M tokens

Llama Prompt Guard 2 86M (Groq)
$0.04 / $0.04 per 1M tokens
Groq
Ultra-fast inference for open models with transparent per-token pricing.
There are 7 published models here. On a 70/30 mix, combined list rates run from $0.03 to $1.32 per 1 million tokens, with a median of $0.1425. The largest listed context window in this group is 131,072 tokens. The most recent price check in this group is 2026-08-19.
Groq list rates
USD per million tokens. The 70/30 mix is a planning default — change it on the model page if your traffic is output-heavy.
| Model | Input / 1M | Output / 1M | Mixed 70/30 | Context | Verified |
|---|---|---|---|---|---|
| Llama Prompt Guard 2 22M (Groq) | $0.03 | $0.03 | $0.03 | 512 | |
| Llama Prompt Guard 2 86M (Groq) | $0.04 | $0.04 | $0.04 | 512 | |
| Llama 3.1 8B Instant (Groq) | $0.05 | $0.08 | $0.059 | 131,072 | |
| GPT OSS 20B (Groq) | $0.075 | $0.3 | $0.1425 | 131,072 | |
| GPT OSS 120B (Groq) | $0.15 | $0.6 | $0.285 | 131,072 | |
| Llama 3.3 70B Versatile (Groq) | $0.59 | $0.79 | $0.65 | 131,072 | |
| Qwen 3.6 27B (Groq) | $0.6 | $3.00 | $1.32 | 131,072 |
Questions about Groq rates
Do these Groq figures include discounts and tax?
No. They are public list rates. Commits, credits, regions, tax, and commercial discounts are not in the number — check the source linked on each model.
When is the more expensive Groq tier worth the extra?
Qwen 3.6 27B (Groq) sits near $1.32 per 1M tokens on a 70/30 mix, versus $0.03 on Llama Prompt Guard 2 22M (Groq). Use the high tier when retries or long reasoning actually fail on the cheap one.