API pricing

Hugging Face Pricing

Hugging Face charges from $0.060 per 1 million output tokens, depending on the model. Llama-3.1-8B-Instruct is the cheapest at $0.060/1M output and $0.060/1M input; its current flagship Qwen3.8 2.4T A95B costs $6.25/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of August 31, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.060/1M

Llama-3.1-8B-Instruct

Cheapest output

$0.060/1M

Llama-3.1-8B-Instruct

Longest context

1.0M

DeepSeek V4 Flash

Models priced

69

Avg $2.08/1M output

Pricing by model

All prices in USD per 1 million tokens. All 69 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Llama-3.1-8B-Instruct
Tools
$0.060$0.060131K
Qwen2.5-Coder-32B-Instruct
Tools
$0.060$0.200131K
Qwen3.5 9B
VisionReasoningTools
$0.170$0.250262K
Qwen3-Coder 30B-A3B Instruct
Tools
$0.070$0.260262K
DeepSeek V4 Flash
ReasoningTools
$0.140$0.2801.0M
DeepSeek V4 Flash 0731
ReasoningTools
$0.140$0.2801.0M
MiMo-V2-Flash
ReasoningTools
$0.100$0.300262K
Step 3.5 Flash
ReasoningTools
$0.100$0.300262K
DeepSeek-V3.2
ReasoningTools
$0.280$0.400164K
Gemma 4 26B A4B IT
VisionReasoningTools
$0.130$0.400262K
Gemma 4 31B IT
VisionReasoningTools
$0.140$0.400262K
GLM-5.3-Flash
VisionReasoningTools
$0.150$0.5001.0M
GPT OSS 20B
ReasoningTools
$0.100$0.500131K
Qwen3 30B A3B
ReasoningTools
$0.120$0.50041K
Hy3
ReasoningTools
$0.140$0.580262K
Qwen3 32B
ReasoningTools
$0.290$0.590131K
GPT OSS 120B
ReasoningTools
$0.250$0.690131K
Llama-3.3-70B-Instruct
Tools
$0.590$0.790131K
Qwen3 235B-A22B
ReasoningTools
$0.200$0.80041K
GLM-4.5-Air
ReasoningTools
$0.130$0.850131K
DeepSeek V4 Pro
ReasoningToolsCache
$0.435$0.8701.0M
GLM-4.6V-Flash
VisionReasoningTools
$0.300$0.900131K
Qwen3.6 35B-A3B
VisionReasoningTools
$0.150$0.950262K
DeepSeek-V3.1
ReasoningTools
$0.270$1.00131K
Qwen3-Next-80B-A3B-Instruct
Tools
$0.250$1.00262K
DeepSeek V3 0324
Tools
$0.270$1.12164K
Step 3.7 Flash
VisionReasoningTools
$0.200$1.15262K
Inkling Small
VisionReasoningTools
$0.500$1.20524K
MiniMax-M2
ReasoningTools
$0.300$1.20205K
MiniMax-M2.1
ReasoningTools
$0.300$1.20205K
MiniMax-M2.5
ReasoningToolsCache
$0.300$1.20205K
MiniMax-M2.7
ReasoningToolsCache
$0.300$1.20205K
MiniMax-M3
VisionReasoningTools
$0.300$1.20524K
DeepSeek-V3
Tools
$0.400$1.3064K
Qwen3 VL 235B A22B Instruct
VisionTools
$0.300$1.50131K
Qwen3-Coder-Next
Tools
$0.200$1.50262K
GLM-4.5V
VisionReasoningTools
$0.600$1.8066K
MiMo-V2.5
ReasoningTools
$0.400$2.00262K
Qwen3-Coder-480B-A35B-Instruct
Tools
$2.00$2.00262K
Qwen3-Next-80B-A3B-Thinking
Tools
$0.300$2.00262K
Qwen3.5 35B-A3B
VisionReasoningTools
$0.250$2.00262K
GLM-4.5
ReasoningTools
$0.600$2.20131K
GLM-4.6
ReasoningTools
$0.550$2.20205K
GLM-4.7
ReasoningToolsCache
$0.600$2.20205K
Qwen3.5 27B
VisionReasoningTools
$0.300$2.40262K
DeepSeek-R1
ReasoningTools
$0.700$2.5064K
Kimi-K2-Thinking
ReasoningToolsCache
$0.600$2.50262K
Qwen3 235B-A22B Instruct 2507
Tools
$0.855$2.56262K
Kimi-K2-Instruct
Tools
$1.00$3.00131K
Kimi-K2-Instruct-0905
Tools
$1.00$3.00262K
Kimi-K2.5
VisionReasoningToolsCache
$0.600$3.00262K
MiMo-V2.5-Pro
ReasoningTools
$1.00$3.001.0M
Qwen3-235B-A22B-Thinking-2507
ReasoningTools
$0.300$3.00262K
Qwen3.8 27B
VisionReasoningTools
$0.400$3.00262K
Qwen3.6 27B
VisionReasoningTools
$0.470$3.19262K
GLM-5
ReasoningToolsCache
$1.00$3.20203K
GLM-5.1
ReasoningToolsCache
$1.00$3.20203K
Qwen3.5 122B-A10B
VisionReasoningTools
$0.400$3.20262K
Qwen3.5-397B-A17B
VisionReasoningTools
$0.600$3.60262K
Qwen3 VL 235B A22B Thinking
VisionReasoningTools
$0.980$3.95131K
DeepSeek V4 Pro 0813
ReasoningTools
$1.32$3.961.0M
Kimi K2.7 Code
VisionReasoningTools
$0.950$4.00262K
Kimi-K2.6
VisionReasoningToolsCache
$0.950$4.00262K
Inkling
VisionReasoningTools
$1.00$4.051.0M
GLM-5.2
ReasoningTools
$1.40$4.40262K
GLM-5.3
ReasoningTools
$1.40$4.401.0M
DeepSeek-R1-0528
ReasoningTools
$3.00$5.00164K
Qwen3.8 2.4T A95B
ReasoningTools
$2.50$6.25262K
Kimi K3
VisionReasoningTools
$3.00$15.001.0M

Frequently asked questions

How much does Hugging Face cost per 1M tokens?

Hugging Face pricing starts at $0.060 per 1 million output tokens (Llama-3.1-8B-Instruct) across 69 models, with its current flagship Qwen3.8 2.4T A95B at $6.25. Older premium models in the catalog list higher. Rates current as of August 31, 2026.

What is the cheapest Hugging Face model?

Llama-3.1-8B-Instruct is the cheapest Hugging Face model on both axes — $0.060 per 1M input tokens and $0.060 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Hugging Face context window?

DeepSeek V4 Flash has the largest context window in the Hugging Face catalog at 1.0M input tokens, with up to 384K output tokens per response.

Does Hugging Face support prompt caching?

Yes. DeepSeek V4 Pro reads cached input at $0.0036 per 1M tokens, against $0.435 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Hugging Face models support reasoning or vision?

The Hugging Face catalog includes 55 reasoning models, 23 vision models, 69 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Hugging Face bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator