API pricing

Groq Pricing

Groq charges between $0.030 and $0.600 per 1 million output tokens, depending on the model. Llama Prompt Guard 2 22M is the cheapest at $0.030/1M output and $0.030/1M input; GPT OSS 120B is the most expensive at $0.600/1M output. The largest context window is 262K tokens (Moonshotai Kimi K2 Instruct 0905). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.030/1M

Llama Prompt Guard 2 22M

Cheapest output

$0.030/1M

Llama Prompt Guard 2 22M

Longest context

262K

Moonshotai Kimi K2 Instruct 0905

Models priced

20

Avg $0.498/1M output

Pricing by model

All prices in USD per 1 million tokens. All 20 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Llama Prompt Guard 2 22M$0.030$0.030512
Prompt Guard 2 86M$0.040$0.040512
Gemma 7B It$0.050$0.0808K
Llama 3.1 8B
Tools
$0.050$0.080131K
Llama 3.1 8B Instant$0.050$0.080128K
Meta Llama Llama Guard 4 12B$0.200$0.2008K
GPT OSS 20B
ReasoningToolsCache
$0.075$0.300131K
OpenAI GPT Oss 20B$0.075$0.300131K
OpenAI GPT Oss Safeguard 20B$0.075$0.300131K
Safety GPT OSS 20B
ReasoningTools
$0.075$0.300131K
Llama 4 Scout 17B 16E
VisionTools
$0.110$0.340131K
Meta Llama Llama 4 Scout 17B 16e Instruct$0.110$0.340131K
Qwen Qwen3 32B$0.290$0.590131K
Qwen3-32B
ReasoningTools
$0.290$0.590131K
GPT OSS 120B
ReasoningToolsCache
$0.150$0.600131K
Meta Llama Llama 4 Maverick 17B 128e Instruct$0.200$0.600131K
OpenAI GPT Oss 120B$0.150$0.600131K
Llama 3.3 70B
Tools
$0.590$0.790131K
Llama 3.3 70B Versatile$0.590$0.790128K
Moonshotai Kimi K2 Instruct 0905$1.00$3.00262K

Frequently asked questions

How much does Groq cost per 1M tokens?

Groq pricing ranges from $0.030 to $0.600 per 1 million output tokens across 20 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Groq model?

Llama Prompt Guard 2 22M is the cheapest Groq model on both axes — $0.030 per 1M input tokens and $0.030 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Groq context window?

Moonshotai Kimi K2 Instruct 0905 has the largest context window in the Groq catalog at 262K input tokens, with up to 16K output tokens per response.

Does Groq support prompt caching?

Yes. GPT OSS 20B reads cached input at $0.037 per 1M tokens, against $0.075 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Groq models support reasoning or vision?

The Groq catalog includes 4 reasoning models, 1 vision models, 7 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Groq bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator