Groq Pricing
Groq charges between $0.030 and $0.600 per 1 million output tokens, depending on the model. Llama Prompt Guard 2 22M is the cheapest at $0.030/1M output and $0.030/1M input; GPT OSS 120B is the most expensive at $0.600/1M output. The largest context window is 262K tokens (Moonshotai Kimi K2 Instruct 0905). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.030/1M
Llama Prompt Guard 2 22M
Cheapest output
$0.030/1M
Llama Prompt Guard 2 22M
Longest context
262K
Moonshotai Kimi K2 Instruct 0905
Models priced
20
Avg $0.498/1M output
Pricing by model
All prices in USD per 1 million tokens. All 20 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Llama Prompt Guard 2 22M | $0.030 | $0.030 | 512 |
| Prompt Guard 2 86M | $0.040 | $0.040 | 512 |
| Gemma 7B It | $0.050 | $0.080 | 8K |
| Llama 3.1 8B Tools | $0.050 | $0.080 | 131K |
| Llama 3.1 8B Instant | $0.050 | $0.080 | 128K |
| Meta Llama Llama Guard 4 12B | $0.200 | $0.200 | 8K |
| GPT OSS 20B ReasoningToolsCache | $0.075 | $0.300 | 131K |
| OpenAI GPT Oss 20B | $0.075 | $0.300 | 131K |
| OpenAI GPT Oss Safeguard 20B | $0.075 | $0.300 | 131K |
| Safety GPT OSS 20B ReasoningTools | $0.075 | $0.300 | 131K |
| Llama 4 Scout 17B 16E VisionTools | $0.110 | $0.340 | 131K |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.110 | $0.340 | 131K |
| Qwen Qwen3 32B | $0.290 | $0.590 | 131K |
| Qwen3-32B ReasoningTools | $0.290 | $0.590 | 131K |
| GPT OSS 120B ReasoningToolsCache | $0.150 | $0.600 | 131K |
| Meta Llama Llama 4 Maverick 17B 128e Instruct | $0.200 | $0.600 | 131K |
| OpenAI GPT Oss 120B | $0.150 | $0.600 | 131K |
| Llama 3.3 70B Tools | $0.590 | $0.790 | 131K |
| Llama 3.3 70B Versatile | $0.590 | $0.790 | 128K |
| Moonshotai Kimi K2 Instruct 0905 | $1.00 | $3.00 | 262K |
Frequently asked questions
How much does Groq cost per 1M tokens?
Groq pricing ranges from $0.030 to $0.600 per 1 million output tokens across 20 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Groq model?
Llama Prompt Guard 2 22M is the cheapest Groq model on both axes — $0.030 per 1M input tokens and $0.030 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Groq context window?
Moonshotai Kimi K2 Instruct 0905 has the largest context window in the Groq catalog at 262K input tokens, with up to 16K output tokens per response.
Does Groq support prompt caching?
Yes. GPT OSS 20B reads cached input at $0.037 per 1M tokens, against $0.075 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Groq models support reasoning or vision?
The Groq catalog includes 4 reasoning models, 1 vision models, 7 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Groq bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator