API pricing

Zhipu AI Pricing

Zhipu AI charges between $0.100 and $22.00 per 1 million output tokens, depending on the model. Glm 4 32B 0414 128K is the cheapest at $0.100/1M output and $0.100/1M input; GLM-5V-Turbo is the most expensive at $22.00/1M output. The largest context window is 1.0M tokens (GLM-5.2). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.070/1M

GLM-4.7-FlashX

Cheapest output

$0.100/1M

Glm 4 32B 0414 128K

Longest context

1.0M

GLM-5.2

Models priced

15

Avg $4.22/1M output

Pricing by model

All prices in USD per 1 million tokens. All 15 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Glm 4 32B 0414 128K$0.100$0.100128K
GLM-4.7-FlashX
ReasoningToolsCache
$0.070$0.400200K
GLM-4.6V
VisionReasoningTools
$0.300$0.900128K
GLM-4.5-Air
ReasoningToolsCache
$0.200$1.10131K
GLM-4.5V
VisionReasoningTools
$0.600$1.8064K
GLM-4.5
ReasoningToolsCache
$0.600$2.20131K
GLM-4.6
ReasoningToolsCache
$0.600$2.20205K
GLM-4.7
ReasoningToolsCache
$0.600$2.20205K
GLM-5
ReasoningToolsCache
$1.00$3.20205K
GLM-5.1
ReasoningToolsCache
$1.40$4.40200K
GLM-5.2
ReasoningToolsCache
$1.40$4.401.0M
Glm 4.5 Airx$1.10$4.50128K
Glm 5 Code$1.20$5.00200K
Glm 4.5 X$2.20$8.90128K
GLM-5V-Turbo
VisionReasoningToolsCache
$5.00$22.00200K

Frequently asked questions

How much does Zhipu AI cost per 1M tokens?

Zhipu AI pricing ranges from $0.100 to $22.00 per 1 million output tokens across 15 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Zhipu AI model?

Glm 4 32B 0414 128K is the cheapest Zhipu AI model on output tokens at $0.100 per 1M, while GLM-4.7-FlashX is cheapest on input at $0.070 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Zhipu AI context window?

GLM-5.2 has the largest context window in the Zhipu AI catalog at 1.0M input tokens, with up to 131K output tokens per response.

Does Zhipu AI support prompt caching?

Yes. GLM-4.7-FlashX reads cached input at $0.010 per 1M tokens, against $0.070 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Zhipu AI models support reasoning or vision?

The Zhipu AI catalog includes 11 reasoning models, 3 vision models, 11 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Zhipu AI bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator