Cheaper overall
GLM 4.5 Air
58% lower blended rate
Higher benchmark scores
GLM 4.5
leads on 13 of 15 figures
Bigger context window
GLM 4.5 Air
131,072 vs 128,000 tokens
Cheaper cached input
GLM 4.5 Air
$0.03 vs $0.11, 73% less
Where each one wins
GLM 4.5
- Higher graduate-level science (GPQA Diamond) 78.2 vs 73.3
- Higher expert-exam performance (Humanity's Last Exam) 13 vs 7
- Higher instruction following (IFBench) 44.1 vs 37.6
- Ahead on 10 more benchmark figures
GLM 4.5 Air
- Cheaper input tokens $0.20 vs $0.60, 67% less
- Cheaper output tokens $1.10 vs $2.20, 50% less
- Cheaper cached input $0.03 vs $0.11, 73% less
- Cheaper on all four workloads
- Larger context window 131,072 tokens
- Higher support-agent performance (ฯยฒ-Bench Telecom) 46.5 vs 43
- Higher Dataviz Elo 1,197 vs 1,169
What four workloads cost
One run, and the same run a thousand times.
GLM 4.5 Air is cheaper on all four workloads, by much the same margin each time โ about 2.6 times, from $0.0008 against $0.0017 on a chat turn to $0.0414 against $0.1208 on a long document.
| Workload | GLM 4.5 | GLM 4.5 Air | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0017 | $0.0008 | $1.70 / $0.7500 | GLM 4.5 Air is 56% cheaper |
| RAG answer 10,000 in / 800 out | $0.0078 | $0.0029 | $7.76 / $2.88 | GLM 4.5 Air is 63% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1208 | $0.0414 | $120.80 / $41.40 | GLM 4.5 Air is 66% cheaper |
| Long document 400,000 in / 3,000 out | GLM 4.5's context window holds 128,000 tokens, so a 400,000-token prompt does not fit. | |||
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | GLM 4.5 | GLM 4.5 Air | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.60 | $0.20 | GLM 4.5 Air is 67% cheaper |
| Any prompt length | Cached input | $0.11 | $0.03 | GLM 4.5 Air is 73% cheaper |
| Any prompt length | Output | $2.20 | $1.10 | GLM 4.5 Air is 50% cheaper |
Blended: GLM 4.5 $1.00, GLM 4.5 Air $0.425 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Both priced by Z.ai. Standard tier, pay-as-you-go.
Benchmarks
They share 16 figures: GLM 4.5 leads on 13, GLM 4.5 Air on 2. The widest gap is 51 points, on Gamedev Elo.
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GLM 4.5 ahead by 4.9
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GLM 4.5 ahead by 6
Instruction following
IFBench โ Precise following of detailed instructions
GLM 4.5 ahead by 6.5
Support-agent performance
ฯยฒ-Bench Telecom โ Tool-using agent tasks in a telecom support setting
GLM 4.5 Air ahead by 3.5
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GLM 4.5 ahead by 6
Command-line work
Terminal-Bench Hard โ Complex command-line and terminal workflows
GLM 4.5 ahead by 1.5
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
Level
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GLM 4.5 ahead by 8.8
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GLM 4.5 ahead by 22.8
3d Elo
Design Arena head-to-head rating in the 3d category
GLM 4.5 ahead by 48
Codecategories Elo
Design Arena head-to-head rating in the codecategories category
GLM 4.5 ahead by 27
Dataviz Elo
Design Arena head-to-head rating in the dataviz category
GLM 4.5 Air ahead by 28
Gamedev Elo
Design Arena head-to-head rating in the gamedev category
GLM 4.5 ahead by 51
Svg Elo
Design Arena head-to-head rating in the svg category
GLM 4.5 ahead by 25
Uicomponent Elo
Design Arena head-to-head rating in the uicomponent category
GLM 4.5 ahead by 21
Website Elo
Design Arena head-to-head rating in the website category
GLM 4.5 ahead by 24
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GLM 4.5 Air holds 3,072 more input tokens in one request. Both list the same capabilities, so nothing here separates them.
| Specification | GLM 4.5 | GLM 4.5 Air |
|---|---|---|
| Context window | 128,000 tokens | 131,072 tokens |
| Max output | 98,304 tokens | 98,304 tokens |
| Takes in | Text | Text |
| Puts out | Text | Text |
| Knowledge cutoff | 31 Dec 2024 | 31 Dec 2024 |
| Released | 26 Jul 2025 | 26 Jul 2025 |
| Status | Active | Active |
| Sold by | Z.ai | Z.ai |
| Tool use | Yes | Yes |
| Structured outputs | Not listed | Not listed |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Not listed | Not listed |
| Computer use | Not listed | Not listed |
| Open weights | Yes | Yes |
| Reasoning | Yes | Yes |
FAQs
- Is GLM 4.5 cheaper than GLM 4.5 Air?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0017 on GLM 4.5 and $0.0008 on GLM 4.5 Air, and the same model is cheaper on every workload on this page.
- How much do GLM 4.5 and GLM 4.5 Air cost per 1M tokens?
- GLM 4.5 costs $0.60 for input and $2.20 for output. GLM 4.5 Air costs $0.20 and $1.10. Standard pay-as-you-go rates, as of 26 Jul 2025.
- Which is better, GLM 4.5 or GLM 4.5 Air?
- On published benchmarks, GLM 4.5. It leads on 13 of the 16 figures both models report, including graduate-level science (GPQA Diamond), where it scores 78.2 against 73.3.
- Does GLM 4.5 or GLM 4.5 Air have a bigger context window?
- GLM 4.5 Air, at 131,072 tokens against 128,000. That is 3,072 more input tokens in a single request.