GLM API pricing, every Z.ai model per million tokens
On 25 September 2026, Z.ai models on this list cost from $0.045 per million input tokens (GLM 5.3 Flash) to $2.80 (GLM 5.3 Prime), and from $0.14 per million output tokens (GLM 5.3 Flash) to $8.80 (GLM 5.3 Prime).
| Model | Input | Output | Cached input | Batch in / out | Context | Your month | Listed |
|---|---|---|---|---|---|---|---|
| GLM 5.3 Prime | $2.80 | $8.80 | $0.56 | 1M | $300.00 | 2026-09-23 | |
| GLM 5.3 FlashX | $0.37 | $1.25 | $0.09 | 1.04858M | $40.95 | 2026-09-18 | |
| GLM 5.3 Flash | $0.045 | $0.14 | $0.01 | $0.06 / $0.20 | 1.31072M | $4.80 | 2026-08-26 |
| GLM 5.3 | $1.40 | $4.40 | $0.26 | $0.45 / $2.00 | 1.31072M | $150.00 | 2026-08-18 |
| GLM 5.2 free tier on OpenRouter | $0.6496 | $2.04 | $0.1206 | 1.04858M | $69.60 | 2026-06-16 | |
| GLM 5.1 | $0.9646 | $3.03 | $0.1791 | 205K | $103.35 | 2026-04-07 | |
| GLM 5V Turbo | $1.20 | $4.00 | $0.24 | 203K | $132.00 | 2026-04-01 | |
| GLM 5 Turbo | $1.20 | $4.00 | $0.24 | 203K | $132.00 | 2026-03-15 | |
| GLM 5 | $0.60 | $1.92 | $0.12 | 205K | $64.80 | 2026-02-11 | |
| GLM 4.7 Flash | $0.0605 | $0.40 | 200K | $9.63 | 2026-01-19 | ||
| GLM 4.7 | $0.60 | $2.20 | $0.11 | 205K | $69.00 | 2025-12-22 | |
| GLM 4.6V | $0.30 | $0.90 | $0.055 | 131K | $31.50 | 2025-12-08 | |
| GLM 4.6 | $0.43 | $1.75 | $0.08 | 205K | $52.05 | 2025-09-30 | |
| GLM 4.5V | $0.60 | $1.80 | $0.11 | 66K | $63.00 | 2025-08-11 | |
| GLM 4.5 | $0.60 | $2.20 | $0.11 | 131K | $69.00 | 2025-07-25 | |
| GLM 4.5 Air | $0.13 | $0.85 | $0.025 | 131K | $20.55 | 2025-07-25 |
Compare Z.ai with every other provider
Common questions
How much does the GLM API cost?
On 25 September 2026, Z.ai models on this list cost from $0.045 per million input tokens (GLM 5.3 Flash) to $2.80 (GLM 5.3 Prime), and from $0.14 per million output tokens (GLM 5.3 Flash) to $8.80 (GLM 5.3 Prime).
What is the cheapest GLM model?
For a typical mix of three input tokens to every output token, the cheapest is GLM 5.3 Flash at $0.045 per million input tokens and $0.14 per million output tokens.
What is the newest GLM model on the list?
GLM 5.3 Prime, listed on 23 September 2026, at $2.80 per million input tokens and $8.80 per million output tokens.
Which GLM model has the largest context window?
GLM 5.3 Flash, with 1.31072M tokens of context.
Does Z.ai offer batch pricing?
Yes, 2 of its models on this list show a batch price, with batch input at a median of 83 percent of the standard price. Batch suits work that can wait for its answer.
Choosing a model for a production agent
Price is one input. Quality on your own cases, latency and the cost of retries decide the rest. Tell me what you are building.