OpenRouter Pricing
OpenRouter charges from $0.030 per 1 million output tokens, depending on the model. Mistral Nemo is the cheapest at $0.030/1M output and $0.019/1M input; its current flagship Hy4 preview costs $2.50/1M output. The largest context window is 2.0M tokens (Grok 4.20). Prices are USD, current as of August 31, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.017/1M
Granite 4.0 Micro
Cheapest output
$0.030/1M
Mistral Nemo
Longest context
2.0M
Grok 4.20
Models priced
420
Avg $9.69/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 420 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Mistral Nemo Tools | $0.019 | $0.030 | 131K |
| Llama 3 8B Lunaris | $0.040 | $0.050 | 8K |
| MythoMax 13B | $0.060 | $0.060 | 8K |
| Ling-3.0-flash ReasoningToolsCache | $0.021 | $0.063 | 262K |
| Llama-3.1-8B-Instruct ToolsCache | $0.050 | $0.080 | 131K |
| Mistral Small 3 | $0.050 | $0.080 | 33K |
| Gemma 3 4B Vision | $0.050 | $0.100 | 131K |
| Granite 4.1 8B ToolsCache | $0.050 | $0.100 | 131K |
| Ministral 3 3B 2512 VisionToolsCache | $0.100 | $0.100 | 131K |
| Mistralai Ministral 3B 2512 | $0.100 | $0.100 | 131K |
| Nex-N2-Mini VisionReasoningToolsCache | $0.025 | $0.100 | 262K |
| OpenAI GPT Oss 20B | $0.020 | $0.100 | 131K |
| Qwen Qwen3 235B A22b 2507 | $0.071 | $0.100 | 262K |
| Reka Edge VisionTools | $0.100 | $0.100 | 16K |
| Granite 4.0 Micro | $0.017 | $0.112 | 131K |
| Laguna XS 2.1 ReasoningToolsCache | $0.060 | $0.120 | 262K |
| Solar Pro 4 ReasoningToolsCache | $0.030 | $0.120 | 524K |
| GPT OSS 20B ReasoningToolsCache | $0.030 | $0.130 | 131K |
| Mistralai Mistral 7B Instruct | $0.130 | $0.130 | 33K |
| Qwen3.7 Flash VisionReasoningToolsCache | $0.030 | $0.130 | 1.0M |
| Nova Micro 1.0 Tools | $0.035 | $0.140 | 128K |
| Phi 4 | $0.070 | $0.140 | 16K |
| Command R7B | $0.037 | $0.150 | 128K |
| Gemma 3 12B VisionTools | $0.050 | $0.150 | 131K |
| Ministral 3 8B 2512 VisionToolsCache | $0.150 | $0.150 | 262K |
| Mistralai Ministral 8B 2512 | $0.150 | $0.150 | 262K |
| Qwen3.5 9B VisionReasoningTools | $0.100 | $0.150 | 262K |
| DeepSeek V4 Flash Latest ReasoningToolsCache | $0.030 | $0.160 | 1.3M |
| GPT OSS 120B ReasoningTools | $0.037 | $0.170 | 131K |
| Hy-MT2-1.8B | $0.044 | $0.177 | 8K |
| DeepSeek V4 Flash ReasoningToolsCache | $0.089 | $0.177 | 1.0M |
| DeepSeek V4 Flash 0731 ReasoningToolsCache | $0.065 | $0.180 | 1.3M |
| Laguna S 2.1 ReasoningToolsCache | $0.090 | $0.180 | 1.0M |
| Llama Guard 4 12B Vision | $0.180 | $0.180 | 164K |
| Qwen Qwen 2.5 Coder 32B Instruct | $0.180 | $0.180 | 34K |
| Qwen3 30B A3B Instruct 2507 Tools | $0.048 | $0.193 | 262K |
| Bytedance Ui Tars 1.5 7B | $0.100 | $0.200 | 131K |
| Ministral 3 14B 2512 VisionToolsCache | $0.200 | $0.200 | 262K |
| Mistral Small 3.2 24B VisionTools | $0.075 | $0.200 | 131K |
| Mistralai Ministral 14B 2512 | $0.200 | $0.200 | 262K |
| Muse Spark 1.2 Contributor VisionReasoningToolsCache | $0.100 | $0.200 | 1.0M |
| Nemotron 3 Nano 30B A3B ReasoningToolsCache | $0.050 | $0.200 | 262K |
| Nemotron 3.5 Lightning 30B A3B ReasoningToolsCache | $0.080 | $0.200 | 262K |
| Nvidia Nemotron 3.5 Lightning | $0.050 | $0.200 | 262K |
| Qwen2.5 7B Instruct Tools | $0.100 | $0.200 | 33K |
| Reka Flash 3 Reasoning | $0.100 | $0.200 | 66K |
| UI-TARS 7B VisionCache | $0.100 | $0.200 | 128K |
| Llama 3.2 1B Instruct | $0.027 | $0.201 | 60K |
| Nova Lite 1.0 VisionTools | $0.060 | $0.240 | 300K |
| Qwen3 14B ReasoningTools | $0.120 | $0.240 | 131K |
| GLM-5.3-Flash VisionReasoningToolsCache | $0.075 | $0.250 | 1.3M |
| Qwen3.5-Flash VisionReasoningTools | $0.065 | $0.260 | 1.0M |
| DeepSeek DeepSeek Chat | $0.140 | $0.280 | 66K |
| DeepSeek DeepSeek Chat V3 0324 | $0.140 | $0.280 | 66K |
| MiMo-V2.5 VisionReasoningToolsCache | $0.140 | $0.280 | 1.1M |
| Qwen3 32B ReasoningTools | $0.080 | $0.280 | 131K |
| Qwen3-Coder 30B-A3B Instruct Tools | $0.070 | $0.280 | 262K |
| Hy-MT2-30B-A3B | $0.074 | $0.295 | 8K |
| Hy-MT2-7B | $0.074 | $0.295 | 8K |
| gpt-oss-safeguard-20b ReasoningToolsCache | $0.075 | $0.300 | 131K |
| Mistralai Mistral Small 3.1 24B Instruct | $0.100 | $0.300 | 131K |
| Mistralai Mistral Small 3.2 24B Instruct | $0.100 | $0.300 | 128K |
| Seed 1.6 Flash VisionReasoningTools | $0.075 | $0.300 | 262K |
| Step 3.5 Flash ReasoningTools | $0.100 | $0.300 | 262K |
| Voxtral Small 24B 2507 ToolsCache | $0.100 | $0.300 | 33K |
| Xiaomi Mimo V2 Flash | $0.100 | $0.300 | 262K |
| Llama 3.2 3B Instruct | $0.050 | $0.330 | 131K |
| Gemma 4 26B A4B IT VisionReasoningTools | $0.070 | $0.340 | 262K |
| Gemma 4 31B IT VisionReasoningToolsCache | $0.090 | $0.340 | 262K |
| Llama 4 Scout VisionToolsCache | $0.110 | $0.340 | 1.3M |
| Qwen3 235B A22B Instruct 2507 ToolsCache | $0.087 | $0.350 | 262K |
| DeepSeek DeepSeek V3.2 | $0.280 | $0.400 | 164K |
| DeepSeek DeepSeek V3.2 Exp | $0.200 | $0.400 | 164K |
| DeepSeek V3.2 ReasoningToolsCache | $0.269 | $0.400 | 164K |
| Gemini 2.5 Flash-Lite VisionReasoningToolsCache | $0.100 | $0.400 | 1.0M |
| GLM-4.7-Flash ReasoningToolsCache | $0.060 | $0.400 | 203K |
| Google Gemini 2.0 Flash 001 | $0.100 | $0.400 | 1.0M |
| GPT-4.1 nano VisionToolsCache | $0.100 | $0.400 | 1.0M |
| GPT-5 Nano VisionReasoningToolsCache | $0.050 | $0.400 | 400K |
| o1-pro VisionReasoning | $150.00 | $600.00 | 200K |
Frequently asked questions
How much does OpenRouter cost per 1M tokens?
OpenRouter pricing starts at $0.030 per 1 million output tokens (Mistral Nemo) across 420 models, with its current flagship Hy4 preview at $2.50. Older premium models in the catalog list higher. Rates current as of August 31, 2026.
What is the cheapest OpenRouter model?
Mistral Nemo is the cheapest OpenRouter model on output tokens at $0.030 per 1M, while Granite 4.0 Micro is cheapest on input at $0.017 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest OpenRouter context window?
Grok 4.20 has the largest context window in the OpenRouter catalog at 2.0M input tokens, with up to 1.8M output tokens per response.
Does OpenRouter support prompt caching?
Yes. Muse Spark 1.2 Contributor reads cached input at $0.0020 per 1M tokens, against $0.100 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Does OpenRouter offer batch pricing?
Yes — 1 OpenRouter model in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Google Gemini 3 Pro Preview bills $6.00 per 1M output in batch against $12.00 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.
Which OpenRouter models support reasoning or vision?
The OpenRouter catalog includes 210 reasoning models, 180 vision models, 270 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual OpenRouter bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator