Cheaper overall
GLM 4.5 Air
75% lower blended rate
Higher benchmark scores
GPT-5.4 Mini
leads on 9 of 9 figures
Bigger context window
GPT-5.4 Mini
400,000 vs 131,072 tokens
Cheaper cached input
GLM 4.5 Air
$0.03 vs $0.075, 60% less
Where each one wins
GPT-5.4 Mini
- Larger context window 400,000 tokens
- Longer maximum output 128,000 tokens
- Higher graduate-level science (GPQA Diamond) 87.5 vs 73.3
- Higher expert-exam performance (Humanity's Last Exam) 28.1 vs 7
- Higher instruction following (IFBench) 73.3 vs 37.6
- Ahead on 6 more benchmark figures
- Offers Batch and Flex pricing not listed for GLM 4.5 Air
GLM 4.5 Air
- Cheaper input tokens $0.20 vs $0.75, 73% less
- Cheaper output tokens $1.10 vs $4.50, 76% less
- Cheaper cached input $0.03 vs $0.075, 60% less
- Cheaper on all four workloads
What four workloads cost
One run, and the same run a thousand times.
GLM 4.5 Air is cheaper on all four workloads, by much the same margin each time โ about 3.7 times, from $0.0008 against $0.0030 on a chat turn to $0.0414 against $0.1380 on a long document.
| Workload | GPT-5.4 Mini | GLM 4.5 Air | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0030 | $0.0008 | $3.00 / $0.7500 | GLM 4.5 Air is 75% cheaper |
| RAG answer 10,000 in / 800 out | $0.0111 | $0.0029 | $11.10 / $2.88 | GLM 4.5 Air is 74% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1380 | $0.0414 | $138.00 / $41.40 | GLM 4.5 Air is 70% cheaper |
| Long document 400,000 in / 3,000 out | GLM 4.5 Air's context window holds 131,072 tokens, so a 400,000-token prompt does not fit. | |||
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | GPT-5.4 Mini | GLM 4.5 Air | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.75 | $0.20 | GLM 4.5 Air is 73% cheaper |
| Any prompt length | Cached input | $0.075 | $0.03 | GLM 4.5 Air is 60% cheaper |
| Any prompt length | Output | $4.50 | $1.10 | GLM 4.5 Air is 76% cheaper |
Blended: GPT-5.4 Mini $1.6875, GLM 4.5 Air $0.425 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
GPT-5.4 Mini priced by OpenAI, GLM 4.5 Air by Z.ai. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.4 Mini lists Batch and Flex, at up to 50% off its own standard rate; GLM 4.5 Air offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GPT-5.4 Mini | GLM 4.5 Air | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.4 Mini | $0.375 in / $2.25 out 50% off Standard | Not offered | — |
| Flex only on GPT-5.4 Mini | $0.375 in / $2.25 out 50% off Standard | Not offered | — |
Benchmarks
GPT-5.4 Mini leads on all 9 figures, furthest ahead on support-agent performance (ฯยฒ-Bench Telecom), by 36.8 points.
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GPT-5.4 Mini ahead by 14.2
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GPT-5.4 Mini ahead by 21.1
Instruction following
IFBench โ Precise following of detailed instructions
GPT-5.4 Mini ahead by 35.7
Support-agent performance
ฯยฒ-Bench Telecom โ Tool-using agent tasks in a telecom support setting
GPT-5.4 Mini ahead by 36.8
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.4 Mini ahead by 30.3
Command-line work
Terminal-Bench Hard โ Complex command-line and terminal workflows
GPT-5.4 Mini ahead by 31.8
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.4 Mini ahead by 10
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GPT-5.4 Mini ahead by 21.2
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GPT-5.4 Mini ahead by 3.1
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.4 Mini holds 268,928 more input tokens in one request. GPT-5.4 Mini also takes file and image.
| Specification | GPT-5.4 Mini | GLM 4.5 Air |
|---|---|---|
| Context window | 400,000 tokens | 131,072 tokens |
| Max output | 128,000 tokens | 98,304 tokens |
| Takes in | Text, image, file | Text |
| Puts out | Text | Text |
| Knowledge cutoff | 31 Aug 2025 | 31 Dec 2024 |
| Released | 17 Mar 2026 | 26 Jul 2025 |
| Status | Active | Active |
| Sold by | OpenAI, Perplexity | Z.ai |
FAQs
- Is GPT-5.4 Mini cheaper than GLM 4.5 Air?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0030 on GPT-5.4 Mini and $0.0008 on GLM 4.5 Air, and the same model is cheaper on every workload on this page.
- How much do GPT-5.4 Mini and GLM 4.5 Air cost per 1M tokens?
- GPT-5.4 Mini costs $0.75 for input and $4.50 for output. GLM 4.5 Air costs $0.20 and $1.10. Standard pay-as-you-go rates, as of 26 Jul 2025.
- Which is better, GPT-5.4 Mini or GLM 4.5 Air?
- On published benchmarks, GPT-5.4 Mini. It leads on 9 of the 9 figures both models report, including graduate-level science (GPQA Diamond), where it scores 87.5 against 73.3.
- Does GPT-5.4 Mini or GLM 4.5 Air offer batch pricing?
- GPT-5.4 Mini only. Its batch input costs $0.375, 50% off its standard rate, and GLM 4.5 Air lists no batch tier at all.
- Does GPT-5.4 Mini or GLM 4.5 Air have a bigger context window?
- GPT-5.4 Mini, at 400,000 tokens against 131,072. That is 268,928 more input tokens in a single request.