API pricing

Hugging Face Pricing

Hugging Face charges between $0.250 and $4.40 per 1 million output tokens, depending on the model. Qwen3.5 9B is the cheapest at $0.250/1M output and $0.170/1M input; GLM-5.2 is the most expensive at $4.40/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.070/1M

Qwen3-Coder 30B-A3B Instruct

Cheapest output

$0.250/1M

Qwen3.5 9B

Longest context

1.0M

DeepSeek V4 Flash

Models priced

48

Avg $1.85/1M output

Pricing by model

All prices in USD per 1 million tokens. All 48 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Qwen3.5 9B
VisionReasoningTools
$0.170$0.250262K
Qwen3-Coder 30B-A3B Instruct
Tools
$0.070$0.260262K
DeepSeek V4 Flash
ReasoningTools
$0.140$0.2801.0M
MiMo-V2-Flash
ReasoningTools
$0.100$0.300262K
Step 3.5 Flash
ReasoningTools
$0.100$0.300262K
DeepSeek-V3.2
ReasoningTools
$0.280$0.400164K
Gemma 4 26B A4B IT
VisionReasoningTools
$0.130$0.400262K
Gemma 4 31B IT
VisionReasoningTools
$0.140$0.400262K
GPT OSS 20B
ReasoningTools
$0.100$0.500131K
Qwen3 32B
ReasoningTools
$0.290$0.590131K
GPT OSS 120B
ReasoningTools
$0.250$0.690131K
Llama-3.3-70B-Instruct
Tools
$0.590$0.790131K
Qwen3 235B-A22B
ReasoningTools
$0.200$0.80041K
GLM-4.5-Air
ReasoningTools
$0.130$0.850131K
DeepSeek V4 Pro
ReasoningToolsCache
$0.435$0.8701.0M
Qwen3.6 35B-A3B
VisionReasoningTools
$0.150$0.950262K
Qwen3-Next-80B-A3B-Instruct
Tools
$0.250$1.00262K
Step 3.7 Flash
VisionReasoningTools
$0.200$1.15262K
MiniMax-M2
ReasoningTools
$0.300$1.20205K
MiniMax-M2.1
ReasoningTools
$0.300$1.20205K
MiniMax-M2.5
ReasoningToolsCache
$0.300$1.20205K
MiniMax-M2.7
ReasoningToolsCache
$0.300$1.20205K
MiniMax-M3
VisionReasoningTools
$0.300$1.20524K
Qwen3-Coder-Next
Tools
$0.200$1.50262K
GLM-4.5V
VisionReasoningTools
$0.600$1.8066K
Qwen3-Coder-480B-A35B-Instruct
Tools
$2.00$2.00262K
Qwen3-Next-80B-A3B-Thinking
Tools
$0.300$2.00262K
Qwen3.5 35B-A3B
VisionReasoningTools
$0.250$2.00262K
GLM-4.5
ReasoningTools
$0.600$2.20131K
GLM-4.6
ReasoningTools
$0.550$2.20205K
GLM-4.7
ReasoningToolsCache
$0.600$2.20205K
Qwen3.5 27B
VisionReasoningTools
$0.300$2.40262K
DeepSeek-R1
ReasoningTools
$0.700$2.5064K
Kimi-K2-Thinking
ReasoningToolsCache
$0.600$2.50262K
Kimi-K2-Instruct
Tools
$1.00$3.00131K
Kimi-K2-Instruct-0905
Tools
$1.00$3.00262K
Kimi-K2.5
VisionReasoningToolsCache
$0.600$3.00262K
MiMo-V2.5-Pro
ReasoningTools
$1.00$3.001.0M
Qwen3-235B-A22B-Thinking-2507
ReasoningTools
$0.300$3.00262K
Qwen3.6 27B
VisionReasoningTools
$0.470$3.19262K
GLM-5
ReasoningToolsCache
$1.00$3.20203K
GLM-5.1
ReasoningToolsCache
$1.00$3.20203K
Qwen3.5 122B-A10B
VisionReasoningTools
$0.400$3.20262K
Qwen3.5-397B-A17B
VisionReasoningTools
$0.600$3.60262K
Kimi K2.7 Code
VisionReasoningTools
$0.950$4.00262K
Kimi-K2.6
VisionReasoningToolsCache
$0.950$4.00262K
GLM-5.2
ReasoningTools
$1.40$4.40262K
DeepSeek-R1-0528
ReasoningTools
$3.00$5.00164K

Frequently asked questions

How much does Hugging Face cost per 1M tokens?

Hugging Face pricing ranges from $0.250 to $4.40 per 1 million output tokens across 48 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Hugging Face model?

Qwen3.5 9B is the cheapest Hugging Face model on output tokens at $0.250 per 1M, while Qwen3-Coder 30B-A3B Instruct is cheapest on input at $0.070 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Hugging Face context window?

DeepSeek V4 Flash has the largest context window in the Hugging Face catalog at 1.0M input tokens, with up to 384K output tokens per response.

Does Hugging Face support prompt caching?

Yes. DeepSeek V4 Pro reads cached input at $0.0036 per 1M tokens, against $0.435 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Hugging Face models support reasoning or vision?

The Hugging Face catalog includes 40 reasoning models, 15 vision models, 48 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Hugging Face bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator