Zhipu AI Pricing
Zhipu AI charges between $0.100 and $22.00 per 1 million output tokens, depending on the model. Glm 4 32B 0414 128K is the cheapest at $0.100/1M output and $0.100/1M input; GLM-5V-Turbo is the most expensive at $22.00/1M output. The largest context window is 1.0M tokens (GLM-5.2). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.070/1M
GLM-4.7-FlashX
Cheapest output
$0.100/1M
Glm 4 32B 0414 128K
Longest context
1.0M
GLM-5.2
Models priced
15
Avg $4.22/1M output
Pricing by model
All prices in USD per 1 million tokens. All 15 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Glm 4 32B 0414 128K | $0.100 | $0.100 | 128K |
| GLM-4.7-FlashX ReasoningToolsCache | $0.070 | $0.400 | 200K |
| GLM-4.6V VisionReasoningTools | $0.300 | $0.900 | 128K |
| GLM-4.5-Air ReasoningToolsCache | $0.200 | $1.10 | 131K |
| GLM-4.5V VisionReasoningTools | $0.600 | $1.80 | 64K |
| GLM-4.5 ReasoningToolsCache | $0.600 | $2.20 | 131K |
| GLM-4.6 ReasoningToolsCache | $0.600 | $2.20 | 205K |
| GLM-4.7 ReasoningToolsCache | $0.600 | $2.20 | 205K |
| GLM-5 ReasoningToolsCache | $1.00 | $3.20 | 205K |
| GLM-5.1 ReasoningToolsCache | $1.40 | $4.40 | 200K |
| GLM-5.2 ReasoningToolsCache | $1.40 | $4.40 | 1.0M |
| Glm 4.5 Airx | $1.10 | $4.50 | 128K |
| Glm 5 Code | $1.20 | $5.00 | 200K |
| Glm 4.5 X | $2.20 | $8.90 | 128K |
| GLM-5V-Turbo VisionReasoningToolsCache | $5.00 | $22.00 | 200K |
Frequently asked questions
How much does Zhipu AI cost per 1M tokens?
Zhipu AI pricing ranges from $0.100 to $22.00 per 1 million output tokens across 15 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Zhipu AI model?
Glm 4 32B 0414 128K is the cheapest Zhipu AI model on output tokens at $0.100 per 1M, while GLM-4.7-FlashX is cheapest on input at $0.070 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Zhipu AI context window?
GLM-5.2 has the largest context window in the Zhipu AI catalog at 1.0M input tokens, with up to 131K output tokens per response.
Does Zhipu AI support prompt caching?
Yes. GLM-4.7-FlashX reads cached input at $0.010 per 1M tokens, against $0.070 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Zhipu AI models support reasoning or vision?
The Zhipu AI catalog includes 11 reasoning models, 3 vision models, 11 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Zhipu AI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator