OpenRouter Pricing
OpenRouter charges from $0.030 per 1 million output tokens, depending on the model. Ling-2.6-flash is the cheapest at $0.030/1M output and $0.010/1M input; its current flagship Qwen3.8 Max costs $6.00/1M output. The largest context window is 2.0M tokens (Grok 4.20). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.010/1M
Ling-2.6-flash
Cheapest output
$0.030/1M
Ling-2.6-flash
Longest context
2.0M
Grok 4.20
Models priced
407
Avg $9.99/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 407 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Ling-2.6-flash ToolsCache | $0.010 | $0.030 | 262K |
| Mistral Nemo Tools | $0.019 | $0.030 | 131K |
| Llama 3 8B Lunaris | $0.040 | $0.050 | 8K |
| MythoMax 13B | $0.060 | $0.060 | 8K |
| Ling-3.0-flash ReasoningToolsCache | $0.021 | $0.063 | 262K |
| Llama 3.1 8B Instruct ToolsCache | $0.050 | $0.080 | 131K |
| Mistral Small 3 | $0.050 | $0.080 | 33K |
| Gemma 3 4B Vision | $0.050 | $0.100 | 131K |
| Granite 4.1 8B ToolsCache | $0.050 | $0.100 | 131K |
| Ministral 3 3B 2512 VisionToolsCache | $0.100 | $0.100 | 131K |
| Mistralai Ministral 3B 2512 | $0.100 | $0.100 | 131K |
| Nex-N2-Mini VisionReasoningToolsCache | $0.025 | $0.100 | 262K |
| OpenAI GPT Oss 20B | $0.020 | $0.100 | 131K |
| Qwen Qwen3 235B A22b 2507 | $0.071 | $0.100 | 262K |
| Reka Edge VisionTools | $0.100 | $0.100 | 16K |
| Granite 4.0 Micro | $0.017 | $0.112 | 131K |
| Gemma 3n 4B | $0.060 | $0.120 | 33K |
| Laguna XS 2.1 ReasoningToolsCache | $0.060 | $0.120 | 262K |
| Solar Pro 4 ReasoningToolsCache | $0.030 | $0.120 | 524K |
| GPT OSS 20B ReasoningToolsCache | $0.030 | $0.130 | 131K |
| Mistralai Mistral 7B Instruct | $0.130 | $0.130 | 33K |
| Qwen3.7 Flash VisionReasoningToolsCache | $0.030 | $0.130 | 1.0M |
| Nova Micro 1.0 Tools | $0.035 | $0.140 | 128K |
| Phi 4 | $0.070 | $0.140 | 16K |
| Command R7B | $0.037 | $0.150 | 128K |
| Gemma 3 12B VisionTools | $0.050 | $0.150 | 131K |
| Ministral 3 8B 2512 VisionToolsCache | $0.150 | $0.150 | 262K |
| Mistralai Ministral 8B 2512 | $0.150 | $0.150 | 262K |
| Qwen3.5 9B VisionReasoningTools | $0.100 | $0.150 | 262K |
| DeepSeek V4 Flash Latest ReasoningToolsCache | $0.080 | $0.160 | 1.0M |
| GPT OSS 120B ReasoningTools | $0.037 | $0.170 | 131K |
| DeepSeek V4 Flash 0731 ReasoningToolsCache | $0.080 | $0.180 | 1.0M |
| Laguna S 2.1 ReasoningToolsCache | $0.090 | $0.180 | 1.0M |
| Llama Guard 4 12B Vision | $0.180 | $0.180 | 1.0M |
| Qwen Qwen 2.5 Coder 32B Instruct | $0.180 | $0.180 | 34K |
| Qwen3 30B A3B Instruct 2507 Tools | $0.048 | $0.193 | 262K |
| Bytedance Ui Tars 1.5 7B | $0.100 | $0.200 | 131K |
| Ministral 3 14B 2512 VisionToolsCache | $0.200 | $0.200 | 262K |
| Mistralai Ministral 14B 2512 | $0.200 | $0.200 | 262K |
| Nemotron 3 Nano 30B A3B ReasoningToolsCache | $0.050 | $0.200 | 262K |
| Qwen2.5 7B Instruct Tools | $0.100 | $0.200 | 33K |
| Reka Flash 3 Reasoning | $0.100 | $0.200 | 66K |
| UI-TARS 7B VisionCache | $0.100 | $0.200 | 128K |
| Llama 3.2 1B Instruct | $0.027 | $0.201 | 60K |
| Hy3 preview ReasoningToolsCache | $0.063 | $0.210 | 262K |
| Nova Lite 1.0 VisionTools | $0.060 | $0.240 | 300K |
| Qwen3 14B ReasoningTools | $0.120 | $0.240 | 131K |
| Mistral Small 3.2 24B VisionTools | $0.094 | $0.250 | 256K |
| Qwen3.5-Flash VisionReasoningTools | $0.065 | $0.260 | 1.0M |
| DeepSeek DeepSeek Chat | $0.140 | $0.280 | 66K |
| DeepSeek DeepSeek Chat V3 0324 | $0.140 | $0.280 | 66K |
| DeepSeek V4 Flash ReasoningToolsCache | $0.140 | $0.280 | 1.0M |
| MiMo-V2.5 VisionReasoningToolsCache | $0.140 | $0.280 | 1.1M |
| Qwen3 32B ReasoningTools | $0.080 | $0.280 | 131K |
| Qwen3-Coder 30B-A3B Instruct Tools | $0.070 | $0.280 | 262K |
| gpt-oss-safeguard-20b ReasoningToolsCache | $0.075 | $0.300 | 131K |
| Llama 4 Scout VisionTools | $0.100 | $0.300 | 1.3M |
| Mistralai Mistral Small 3.1 24B Instruct | $0.100 | $0.300 | 131K |
| Mistralai Mistral Small 3.2 24B Instruct | $0.100 | $0.300 | 128K |
| Seed 1.6 Flash VisionReasoningTools | $0.075 | $0.300 | 262K |
| Step 3.5 Flash ReasoningTools | $0.100 | $0.300 | 262K |
| Voxtral Small 24B 2507 ToolsCache | $0.100 | $0.300 | 32K |
| Xiaomi Mimo V2 Flash | $0.100 | $0.300 | 262K |
| Llama-3.3-70B-Instruct Tools | $0.100 | $0.320 | 131K |
| Llama 3.2 3B Instruct | $0.050 | $0.330 | 131K |
| Gemma 4 31B IT VisionReasoningToolsCache | $0.100 | $0.340 | 262K |
| DeepSeek DeepSeek V3.2 | $0.280 | $0.400 | 164K |
| DeepSeek DeepSeek V3.2 Exp | $0.200 | $0.400 | 164K |
| DeepSeek V3.2 ReasoningToolsCache | $0.269 | $0.400 | 164K |
| Gemini 2.5 Flash-Lite VisionReasoningToolsCache | $0.100 | $0.400 | 1.0M |
| Gemma 4 26B A4B IT VisionReasoningToolsCache | $0.120 | $0.400 | 262K |
| GLM-4.7-Flash ReasoningToolsCache | $0.060 | $0.400 | 203K |
| Google Gemini 2.0 Flash 001 | $0.100 | $0.400 | 1.0M |
| GPT-4.1 nano VisionToolsCache | $0.100 | $0.400 | 1.0M |
| GPT-5 Nano VisionReasoningToolsCache | $0.050 | $0.400 | 400K |
| Hermes 4 70B Reasoning | $0.130 | $0.400 | 131K |
| Llama 3.1 70B Instruct Tools | $0.400 | $0.400 | 131K |
| Nemotron 3 Super 120B A12B ReasoningTools | $0.085 | $0.400 | 1.0M |
| OpenAI GPT 4.1 Nano | $0.100 | $0.400 | 1.0M |
| o1-pro VisionReasoning | $150.00 | $600.00 | 200K |
Frequently asked questions
How much does OpenRouter cost per 1M tokens?
OpenRouter pricing starts at $0.030 per 1 million output tokens (Ling-2.6-flash) across 407 models, with its current flagship Qwen3.8 Max at $6.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest OpenRouter model?
Ling-2.6-flash is the cheapest OpenRouter model on both axes — $0.010 per 1M input tokens and $0.030 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest OpenRouter context window?
Grok 4.20 has the largest context window in the OpenRouter catalog at 2.0M input tokens, with up to 2.0M output tokens per response.
Does OpenRouter support prompt caching?
Yes. Ling-2.6-flash reads cached input at $0.0020 per 1M tokens, against $0.010 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Does OpenRouter offer batch pricing?
Yes — 1 OpenRouter model in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Google Gemini 3 Pro Preview bills $6.00 per 1M output in batch against $12.00 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.
Which OpenRouter models support reasoning or vision?
The OpenRouter catalog includes 199 reasoning models, 171 vision models, 259 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual OpenRouter bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator