Hugging Face Pricing
Hugging Face charges from $0.060 per 1 million output tokens, depending on the model. Llama-3.1-8B-Instruct is the cheapest at $0.060/1M output and $0.060/1M input; its current flagship Qwen3.8 2.4T A95B costs $6.25/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of August 31, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.060/1M
Llama-3.1-8B-Instruct
Cheapest output
$0.060/1M
Llama-3.1-8B-Instruct
Longest context
1.0M
DeepSeek V4 Flash
Models priced
69
Avg $2.08/1M output
Pricing by model
All prices in USD per 1 million tokens. All 69 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Llama-3.1-8B-Instruct Tools | $0.060 | $0.060 | 131K |
| Qwen2.5-Coder-32B-Instruct Tools | $0.060 | $0.200 | 131K |
| Qwen3.5 9B VisionReasoningTools | $0.170 | $0.250 | 262K |
| Qwen3-Coder 30B-A3B Instruct Tools | $0.070 | $0.260 | 262K |
| DeepSeek V4 Flash ReasoningTools | $0.140 | $0.280 | 1.0M |
| DeepSeek V4 Flash 0731 ReasoningTools | $0.140 | $0.280 | 1.0M |
| MiMo-V2-Flash ReasoningTools | $0.100 | $0.300 | 262K |
| Step 3.5 Flash ReasoningTools | $0.100 | $0.300 | 262K |
| DeepSeek-V3.2 ReasoningTools | $0.280 | $0.400 | 164K |
| Gemma 4 26B A4B IT VisionReasoningTools | $0.130 | $0.400 | 262K |
| Gemma 4 31B IT VisionReasoningTools | $0.140 | $0.400 | 262K |
| GLM-5.3-Flash VisionReasoningTools | $0.150 | $0.500 | 1.0M |
| GPT OSS 20B ReasoningTools | $0.100 | $0.500 | 131K |
| Qwen3 30B A3B ReasoningTools | $0.120 | $0.500 | 41K |
| Hy3 ReasoningTools | $0.140 | $0.580 | 262K |
| Qwen3 32B ReasoningTools | $0.290 | $0.590 | 131K |
| GPT OSS 120B ReasoningTools | $0.250 | $0.690 | 131K |
| Llama-3.3-70B-Instruct Tools | $0.590 | $0.790 | 131K |
| Qwen3 235B-A22B ReasoningTools | $0.200 | $0.800 | 41K |
| GLM-4.5-Air ReasoningTools | $0.130 | $0.850 | 131K |
| DeepSeek V4 Pro ReasoningToolsCache | $0.435 | $0.870 | 1.0M |
| GLM-4.6V-Flash VisionReasoningTools | $0.300 | $0.900 | 131K |
| Qwen3.6 35B-A3B VisionReasoningTools | $0.150 | $0.950 | 262K |
| DeepSeek-V3.1 ReasoningTools | $0.270 | $1.00 | 131K |
| Qwen3-Next-80B-A3B-Instruct Tools | $0.250 | $1.00 | 262K |
| DeepSeek V3 0324 Tools | $0.270 | $1.12 | 164K |
| Step 3.7 Flash VisionReasoningTools | $0.200 | $1.15 | 262K |
| Inkling Small VisionReasoningTools | $0.500 | $1.20 | 524K |
| MiniMax-M2 ReasoningTools | $0.300 | $1.20 | 205K |
| MiniMax-M2.1 ReasoningTools | $0.300 | $1.20 | 205K |
| MiniMax-M2.5 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax-M2.7 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax-M3 VisionReasoningTools | $0.300 | $1.20 | 524K |
| DeepSeek-V3 Tools | $0.400 | $1.30 | 64K |
| Qwen3 VL 235B A22B Instruct VisionTools | $0.300 | $1.50 | 131K |
| Qwen3-Coder-Next Tools | $0.200 | $1.50 | 262K |
| GLM-4.5V VisionReasoningTools | $0.600 | $1.80 | 66K |
| MiMo-V2.5 ReasoningTools | $0.400 | $2.00 | 262K |
| Qwen3-Coder-480B-A35B-Instruct Tools | $2.00 | $2.00 | 262K |
| Qwen3-Next-80B-A3B-Thinking Tools | $0.300 | $2.00 | 262K |
| Qwen3.5 35B-A3B VisionReasoningTools | $0.250 | $2.00 | 262K |
| GLM-4.5 ReasoningTools | $0.600 | $2.20 | 131K |
| GLM-4.6 ReasoningTools | $0.550 | $2.20 | 205K |
| GLM-4.7 ReasoningToolsCache | $0.600 | $2.20 | 205K |
| Qwen3.5 27B VisionReasoningTools | $0.300 | $2.40 | 262K |
| DeepSeek-R1 ReasoningTools | $0.700 | $2.50 | 64K |
| Kimi-K2-Thinking ReasoningToolsCache | $0.600 | $2.50 | 262K |
| Qwen3 235B-A22B Instruct 2507 Tools | $0.855 | $2.56 | 262K |
| Kimi-K2-Instruct Tools | $1.00 | $3.00 | 131K |
| Kimi-K2-Instruct-0905 Tools | $1.00 | $3.00 | 262K |
| Kimi-K2.5 VisionReasoningToolsCache | $0.600 | $3.00 | 262K |
| MiMo-V2.5-Pro ReasoningTools | $1.00 | $3.00 | 1.0M |
| Qwen3-235B-A22B-Thinking-2507 ReasoningTools | $0.300 | $3.00 | 262K |
| Qwen3.8 27B VisionReasoningTools | $0.400 | $3.00 | 262K |
| Qwen3.6 27B VisionReasoningTools | $0.470 | $3.19 | 262K |
| GLM-5 ReasoningToolsCache | $1.00 | $3.20 | 203K |
| GLM-5.1 ReasoningToolsCache | $1.00 | $3.20 | 203K |
| Qwen3.5 122B-A10B VisionReasoningTools | $0.400 | $3.20 | 262K |
| Qwen3.5-397B-A17B VisionReasoningTools | $0.600 | $3.60 | 262K |
| Qwen3 VL 235B A22B Thinking VisionReasoningTools | $0.980 | $3.95 | 131K |
| DeepSeek V4 Pro 0813 ReasoningTools | $1.32 | $3.96 | 1.0M |
| Kimi K2.7 Code VisionReasoningTools | $0.950 | $4.00 | 262K |
| Kimi-K2.6 VisionReasoningToolsCache | $0.950 | $4.00 | 262K |
| Inkling VisionReasoningTools | $1.00 | $4.05 | 1.0M |
| GLM-5.2 ReasoningTools | $1.40 | $4.40 | 262K |
| GLM-5.3 ReasoningTools | $1.40 | $4.40 | 1.0M |
| DeepSeek-R1-0528 ReasoningTools | $3.00 | $5.00 | 164K |
| Qwen3.8 2.4T A95B ReasoningTools | $2.50 | $6.25 | 262K |
| Kimi K3 VisionReasoningTools | $3.00 | $15.00 | 1.0M |
Frequently asked questions
How much does Hugging Face cost per 1M tokens?
Hugging Face pricing starts at $0.060 per 1 million output tokens (Llama-3.1-8B-Instruct) across 69 models, with its current flagship Qwen3.8 2.4T A95B at $6.25. Older premium models in the catalog list higher. Rates current as of August 31, 2026.
What is the cheapest Hugging Face model?
Llama-3.1-8B-Instruct is the cheapest Hugging Face model on both axes — $0.060 per 1M input tokens and $0.060 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Hugging Face context window?
DeepSeek V4 Flash has the largest context window in the Hugging Face catalog at 1.0M input tokens, with up to 384K output tokens per response.
Does Hugging Face support prompt caching?
Yes. DeepSeek V4 Pro reads cached input at $0.0036 per 1M tokens, against $0.435 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Hugging Face models support reasoning or vision?
The Hugging Face catalog includes 55 reasoning models, 23 vision models, 69 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Hugging Face bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator