API pricing

OpenRouter Pricing

OpenRouter charges from $0.030 per 1 million output tokens, depending on the model. Ling-2.6-flash is the cheapest at $0.030/1M output and $0.010/1M input; its current flagship Qwen3.8 Max costs $6.00/1M output. The largest context window is 2.0M tokens (Grok 4.20). Prices are USD, current as of August 11, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.010/1M

Ling-2.6-flash

Cheapest output

$0.030/1M

Ling-2.6-flash

Longest context

2.0M

Grok 4.20

Models priced

407

Avg $9.99/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 407 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Ling-2.6-flash
ToolsCache
$0.010$0.030262K
Mistral Nemo
Tools
$0.019$0.030131K
Llama 3 8B Lunaris$0.040$0.0508K
MythoMax 13B$0.060$0.0608K
Ling-3.0-flash
ReasoningToolsCache
$0.021$0.063262K
Llama 3.1 8B Instruct
ToolsCache
$0.050$0.080131K
Mistral Small 3$0.050$0.08033K
Gemma 3 4B
Vision
$0.050$0.100131K
Granite 4.1 8B
ToolsCache
$0.050$0.100131K
Ministral 3 3B 2512
VisionToolsCache
$0.100$0.100131K
Mistralai Ministral 3B 2512$0.100$0.100131K
Nex-N2-Mini
VisionReasoningToolsCache
$0.025$0.100262K
OpenAI GPT Oss 20B$0.020$0.100131K
Qwen Qwen3 235B A22b 2507$0.071$0.100262K
Reka Edge
VisionTools
$0.100$0.10016K
Granite 4.0 Micro$0.017$0.112131K
Gemma 3n 4B$0.060$0.12033K
Laguna XS 2.1
ReasoningToolsCache
$0.060$0.120262K
Solar Pro 4
ReasoningToolsCache
$0.030$0.120524K
GPT OSS 20B
ReasoningToolsCache
$0.030$0.130131K
Mistralai Mistral 7B Instruct$0.130$0.13033K
Qwen3.7 Flash
VisionReasoningToolsCache
$0.030$0.1301.0M
Nova Micro 1.0
Tools
$0.035$0.140128K
Phi 4$0.070$0.14016K
Command R7B$0.037$0.150128K
Gemma 3 12B
VisionTools
$0.050$0.150131K
Ministral 3 8B 2512
VisionToolsCache
$0.150$0.150262K
Mistralai Ministral 8B 2512$0.150$0.150262K
Qwen3.5 9B
VisionReasoningTools
$0.100$0.150262K
DeepSeek V4 Flash Latest
ReasoningToolsCache
$0.080$0.1601.0M
GPT OSS 120B
ReasoningTools
$0.037$0.170131K
DeepSeek V4 Flash 0731
ReasoningToolsCache
$0.080$0.1801.0M
Laguna S 2.1
ReasoningToolsCache
$0.090$0.1801.0M
Llama Guard 4 12B
Vision
$0.180$0.1801.0M
Qwen Qwen 2.5 Coder 32B Instruct$0.180$0.18034K
Qwen3 30B A3B Instruct 2507
Tools
$0.048$0.193262K
Bytedance Ui Tars 1.5 7B$0.100$0.200131K
Ministral 3 14B 2512
VisionToolsCache
$0.200$0.200262K
Mistralai Ministral 14B 2512$0.200$0.200262K
Nemotron 3 Nano 30B A3B
ReasoningToolsCache
$0.050$0.200262K
Qwen2.5 7B Instruct
Tools
$0.100$0.20033K
Reka Flash 3
Reasoning
$0.100$0.20066K
UI-TARS 7B
VisionCache
$0.100$0.200128K
Llama 3.2 1B Instruct$0.027$0.20160K
Hy3 preview
ReasoningToolsCache
$0.063$0.210262K
Nova Lite 1.0
VisionTools
$0.060$0.240300K
Qwen3 14B
ReasoningTools
$0.120$0.240131K
Mistral Small 3.2 24B
VisionTools
$0.094$0.250256K
Qwen3.5-Flash
VisionReasoningTools
$0.065$0.2601.0M
DeepSeek DeepSeek Chat$0.140$0.28066K
DeepSeek DeepSeek Chat V3 0324$0.140$0.28066K
DeepSeek V4 Flash
ReasoningToolsCache
$0.140$0.2801.0M
MiMo-V2.5
VisionReasoningToolsCache
$0.140$0.2801.1M
Qwen3 32B
ReasoningTools
$0.080$0.280131K
Qwen3-Coder 30B-A3B Instruct
Tools
$0.070$0.280262K
gpt-oss-safeguard-20b
ReasoningToolsCache
$0.075$0.300131K
Llama 4 Scout
VisionTools
$0.100$0.3001.3M
Mistralai Mistral Small 3.1 24B Instruct$0.100$0.300131K
Mistralai Mistral Small 3.2 24B Instruct$0.100$0.300128K
Seed 1.6 Flash
VisionReasoningTools
$0.075$0.300262K
Step 3.5 Flash
ReasoningTools
$0.100$0.300262K
Voxtral Small 24B 2507
ToolsCache
$0.100$0.30032K
Xiaomi Mimo V2 Flash$0.100$0.300262K
Llama-3.3-70B-Instruct
Tools
$0.100$0.320131K
Llama 3.2 3B Instruct$0.050$0.330131K
Gemma 4 31B IT
VisionReasoningToolsCache
$0.100$0.340262K
DeepSeek DeepSeek V3.2$0.280$0.400164K
DeepSeek DeepSeek V3.2 Exp$0.200$0.400164K
DeepSeek V3.2
ReasoningToolsCache
$0.269$0.400164K
Gemini 2.5 Flash-Lite
VisionReasoningToolsCache
$0.100$0.4001.0M
Gemma 4 26B A4B IT
VisionReasoningToolsCache
$0.120$0.400262K
GLM-4.7-Flash
ReasoningToolsCache
$0.060$0.400203K
Google Gemini 2.0 Flash 001$0.100$0.4001.0M
GPT-4.1 nano
VisionToolsCache
$0.100$0.4001.0M
GPT-5 Nano
VisionReasoningToolsCache
$0.050$0.400400K
Hermes 4 70B
Reasoning
$0.130$0.400131K
Llama 3.1 70B Instruct
Tools
$0.400$0.400131K
Nemotron 3 Super 120B A12B
ReasoningTools
$0.085$0.4001.0M
OpenAI GPT 4.1 Nano$0.100$0.4001.0M
o1-pro
VisionReasoning
$150.00$600.00200K

Frequently asked questions

How much does OpenRouter cost per 1M tokens?

OpenRouter pricing starts at $0.030 per 1 million output tokens (Ling-2.6-flash) across 407 models, with its current flagship Qwen3.8 Max at $6.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.

What is the cheapest OpenRouter model?

Ling-2.6-flash is the cheapest OpenRouter model on both axes — $0.010 per 1M input tokens and $0.030 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest OpenRouter context window?

Grok 4.20 has the largest context window in the OpenRouter catalog at 2.0M input tokens, with up to 2.0M output tokens per response.

Does OpenRouter support prompt caching?

Yes. Ling-2.6-flash reads cached input at $0.0020 per 1M tokens, against $0.010 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Does OpenRouter offer batch pricing?

Yes — 1 OpenRouter model in this catalog publish a discounted rate for asynchronous batch jobs, at 50% off output tokens. Google Gemini 3 Pro Preview bills $6.00 per 1M output in batch against $12.00 synchronously. Batch trades real-time responses for the lower rate, so it suits backfills, evaluations and bulk enrichment rather than user-facing calls.

Which OpenRouter models support reasoning or vision?

The OpenRouter catalog includes 199 reasoning models, 171 vision models, 259 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual OpenRouter bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator