API pricing

OpenRouter Pricing

OpenRouter charges from $0.030 per 1 million output tokens, depending on the model. Mistral Nemo is the cheapest at $0.030/1M output and $0.019/1M input; its current flagship Hy4 preview costs $2.50/1M output. The largest context window is 2.0M tokens (Grok 4.20). Prices are USD, current as of August 31, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.017/1M

Granite 4.0 Micro

Cheapest output

$0.030/1M

Mistral Nemo

Longest context

2.0M

Grok 4.20

Models priced

420

Avg $9.69/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 420 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Mistral Nemo
Tools
$0.019$0.030131K
Llama 3 8B Lunaris$0.040$0.0508K
MythoMax 13B$0.060$0.0608K
Ling-3.0-flash
ReasoningToolsCache
$0.021$0.063262K
Llama-3.1-8B-Instruct
ToolsCache
$0.050$0.080131K
Mistral Small 3$0.050$0.08033K
Gemma 3 4B
Vision
$0.050$0.100131K
Granite 4.1 8B
ToolsCache
$0.050$0.100131K
Ministral 3 3B 2512
VisionToolsCache
$0.100$0.100131K
Mistralai Ministral 3B 2512$0.100$0.100131K
Nex-N2-Mini
VisionReasoningToolsCache
$0.025$0.100262K
OpenAI GPT Oss 20B$0.020$0.100131K
Qwen Qwen3 235B A22b 2507$0.071$0.100262K
Reka Edge
VisionTools
$0.100$0.10016K
Granite 4.0 Micro$0.017$0.112131K
Laguna XS 2.1
ReasoningToolsCache
$0.060$0.120262K
Solar Pro 4
ReasoningToolsCache
$0.030$0.120524K
GPT OSS 20B
ReasoningToolsCache
$0.030$0.130131K
Mistralai Mistral 7B Instruct$0.130$0.13033K
Qwen3.7 Flash
VisionReasoningToolsCache
$0.030$0.1301.0M
Nova Micro 1.0
Tools
$0.035$0.140128K
Phi 4$0.070$0.14016K
Command R7B$0.037$0.150128K
Gemma 3 12B
VisionTools
$0.050$0.150131K
Ministral 3 8B 2512
VisionToolsCache
$0.150$0.150262K
Mistralai Ministral 8B 2512$0.150$0.150262K
Qwen3.5 9B
VisionReasoningTools
$0.100$0.150262K
DeepSeek V4 Flash Latest
ReasoningToolsCache
$0.030$0.1601.3M
GPT OSS 120B
ReasoningTools
$0.037$0.170131K
Hy-MT2-1.8B$0.044$0.1778K
DeepSeek V4 Flash
ReasoningToolsCache
$0.089$0.1771.0M
DeepSeek V4 Flash 0731
ReasoningToolsCache
$0.065$0.1801.3M
Laguna S 2.1
ReasoningToolsCache
$0.090$0.1801.0M
Llama Guard 4 12B
Vision
$0.180$0.180164K
Qwen Qwen 2.5 Coder 32B Instruct$0.180$0.18034K
Qwen3 30B A3B Instruct 2507
Tools
$0.048$0.193262K
Bytedance Ui Tars 1.5 7B$0.100$0.200131K
Ministral 3 14B 2512
VisionToolsCache
$0.200$0.200262K
Mistral Small 3.2 24B
VisionTools
$0.075$0.200131K
Mistralai Ministral 14B 2512$0.200$0.200262K
Muse Spark 1.2 Contributor
VisionReasoningToolsCache
$0.100$0.2001.0M
Nemotron 3 Nano 30B A3B
ReasoningToolsCache
$0.050$0.200262K
Nemotron 3.5 Lightning 30B A3B
ReasoningToolsCache
$0.080$0.200262K
Nvidia Nemotron 3.5 Lightning$0.050$0.200262K
Qwen2.5 7B Instruct
Tools
$0.100$0.20033K
Reka Flash 3
Reasoning
$0.100$0.20066K
UI-TARS 7B
VisionCache
$0.100$0.200128K
Llama 3.2 1B Instruct$0.027$0.20160K
Nova Lite 1.0
VisionTools
$0.060$0.240300K
Qwen3 14B
ReasoningTools
$0.120$0.240131K
GLM-5.3-Flash
VisionReasoningToolsCache
$0.075$0.2501.3M
Qwen3.5-Flash
VisionReasoningTools
$0.065$0.2601.0M
DeepSeek DeepSeek Chat$0.140$0.28066K
DeepSeek DeepSeek Chat V3 0324$0.140$0.28066K
MiMo-V2.5
VisionReasoningToolsCache
$0.140$0.2801.1M
Qwen3 32B
ReasoningTools
$0.080$0.280131K
Qwen3-Coder 30B-A3B Instruct
Tools
$0.070$0.280262K
Hy-MT2-30B-A3B$0.074$0.2958K
Hy-MT2-7B$0.074$0.2958K
gpt-oss-safeguard-20b
ReasoningToolsCache
$0.075$0.300131K
Mistralai Mistral Small 3.1 24B Instruct$0.100$0.300131K
Mistralai Mistral Small 3.2 24B Instruct$0.100$0.300128K
Seed 1.6 Flash
VisionReasoningTools
$0.075$0.300262K
Step 3.5 Flash
ReasoningTools
$0.100$0.300262K
Voxtral Small 24B 2507
ToolsCache
$0.100$0.30033K
Xiaomi Mimo V2 Flash$0.100$0.300262K
Llama 3.2 3B Instruct$0.050$0.330131K
Gemma 4 26B A4B IT
VisionReasoningTools
$0.070$0.340262K
Gemma 4 31B IT
VisionReasoningToolsCache
$0.090$0.340262K
Llama 4 Scout
VisionToolsCache
$0.110$0.3401.3M
Qwen3 235B A22B Instruct 2507
ToolsCache
$0.087$0.350262K
DeepSeek DeepSeek V3.2$0.280$0.400164K
DeepSeek DeepSeek V3.2 Exp$0.200$0.400164K
DeepSeek V3.2
ReasoningToolsCache
$0.269$0.400164K
Gemini 2.5 Flash-Lite
VisionReasoningToolsCache
$0.100$0.4001.0M
GLM-4.7-Flash
ReasoningToolsCache
$0.060$0.400203K
Google Gemini 2.0 Flash 001$0.100$0.4001.0M
GPT-4.1 nano
VisionToolsCache
$0.100$0.4001.0M
GPT-5 Nano
VisionReasoningToolsCache
$0.050$0.400400K
o1-pro
VisionReasoning
$150.00$600.00200K

Frequently asked questions

How much does OpenRouter cost per 1M tokens?

OpenRouter pricing starts at $0.030 per 1 million output tokens (Mistral Nemo) across 420 models, with its current flagship Hy4 preview at $2.50. Older premium models in the catalog list higher. Rates current as of August 31, 2026.

What is the cheapest OpenRouter model?

Mistral Nemo is the cheapest OpenRouter model on output tokens at $0.030 per 1M, while Granite 4.0 Micro is cheapest on input at $0.017 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest OpenRouter context window?

Grok 4.20 has the largest context window in the OpenRouter catalog at 2.0M input tokens, with up to 1.8M output tokens per response.

Does OpenRouter support prompt caching?

Yes. Muse Spark 1.2 Contributor reads cached input at $0.0020 per 1M tokens, against $0.100 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Does OpenRouter offer batch pricing?

Yes — 1 OpenRouter model in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Google Gemini 3 Pro Preview bills $6.00 per 1M output in batch against $12.00 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.

Which OpenRouter models support reasoning or vision?

The OpenRouter catalog includes 210 reasoning models, 180 vision models, 270 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual OpenRouter bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator