Cheaper overall
GLM 4.6
41% lower blended rate
Higher benchmark scores
GPT-5.4 Mini
leads on 10 of 11 figures
Bigger context window
GPT-5.4 Mini
400,000 vs 204,800 tokens
Cheaper cached input
GPT-5.4 Mini
$0.075 vs $0.11, 32% less
Where each one wins
GPT-5.4 Mini
- Cheaper cached input $0.075 vs $0.11, 32% less
- Larger context window 400,000 tokens
- Longer maximum output 128,000 tokens
- Higher coding ability (Coding Index) 56.1 vs 45.8
- Higher graduate-level science (GPQA Diamond) 87.5 vs 78
- Higher expert-exam performance (Humanity's Last Exam) 28.1 vs 14.5
- Ahead on 7 more benchmark figures
- Offers Batch and Flex pricing not listed for GLM 4.6
GLM 4.6
- Cheaper input tokens $0.60 vs $0.75, 20% less
- Cheaper output tokens $2.20 vs $4.50, 51% less
- Cheaper on all four workloads
- Higher answer reliability (Omniscience Non-Hallucination) 32.4 vs 10.2
What four workloads cost
One run, and the same run a thousand times.
GLM 4.6 is cheaper on all four workloads, by much the same margin each time โ about 1.4 times, from $0.0017 against $0.0030 on a chat turn to $0.1208 against $0.1380 on a long document.
| Workload | GPT-5.4 Mini | GLM 4.6 | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0030 | $0.0017 | $3.00 / $1.70 | GLM 4.6 is 43% cheaper |
| RAG answer 10,000 in / 800 out | $0.0111 | $0.0078 | $11.10 / $7.76 | GLM 4.6 is 30% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1380 | $0.1208 | $138.00 / $120.80 | GLM 4.6 is 12% cheaper |
| Long document 400,000 in / 3,000 out | GLM 4.6's context window holds 204,800 tokens, so a 400,000-token prompt does not fit. | |||
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | GPT-5.4 Mini | GLM 4.6 | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.75 | $0.60 | GLM 4.6 is 20% cheaper |
| Any prompt length | Cached input | $0.075 | $0.11 | GPT-5.4 Mini is 32% cheaper |
| Any prompt length | Output | $4.50 | $2.20 | GLM 4.6 is 51% cheaper |
Blended: GPT-5.4 Mini $1.6875, GLM 4.6 $1.00 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
GPT-5.4 Mini priced by OpenAI, GLM 4.6 by Z.ai. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.4 Mini lists Batch and Flex, at up to 50% off its own standard rate; GLM 4.6 offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GPT-5.4 Mini | GLM 4.6 | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.4 Mini | $0.375 in / $2.25 out 50% off Standard | Not offered | — |
| Flex only on GPT-5.4 Mini | $0.375 in / $2.25 out 50% off Standard | Not offered | — |
Benchmarks
They share 11 figures: GPT-5.4 Mini leads on 10, GLM 4.6 on 1. The widest gap is 29.9 points, on instruction following (IFBench).
Coding ability
Coding Index โ Coding ability across software-engineering evals
GPT-5.4 Mini ahead by 10.3
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GPT-5.4 Mini ahead by 9.5
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GPT-5.4 Mini ahead by 13.6
Instruction following
IFBench โ Precise following of detailed instructions
GPT-5.4 Mini ahead by 29.9
Support-agent performance
ฯยฒ-Bench Telecom โ Tool-using agent tasks in a telecom support setting
GPT-5.4 Mini ahead by 6.4
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.4 Mini ahead by 23
Command-line work
Terminal-Bench Hard โ Complex command-line and terminal workflows
GPT-5.4 Mini ahead by 23.5
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.4 Mini ahead by 8.9
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
GPT-5.4 Mini ahead by 12.7
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GPT-5.4 Mini ahead by 10.6
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GLM 4.6 ahead by 22.2
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.4 Mini holds 195,200 more input tokens in one request. GPT-5.4 Mini also takes file and image.
| Specification | GPT-5.4 Mini | GLM 4.6 |
|---|---|---|
| Context window | 400,000 tokens | 204,800 tokens |
| Max output | 128,000 tokens | 16,384 tokens |
| Takes in | Text, image, file | Text |
| Puts out | Text | Text |
| Knowledge cutoff | 31 Aug 2025 | 31 Mar 2025 |
| Released | 17 Mar 2026 | 30 Sep 2025 |
| Status | Active | Active |
| Sold by | OpenAI, Perplexity | Z.ai |
FAQs
- Is GPT-5.4 Mini cheaper than GLM 4.6?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0030 on GPT-5.4 Mini and $0.0017 on GLM 4.6, and the same model is cheaper on every workload on this page.
- How much do GPT-5.4 Mini and GLM 4.6 cost per 1M tokens?
- GPT-5.4 Mini costs $0.75 for input and $4.50 for output. GLM 4.6 costs $0.60 and $2.20. Standard pay-as-you-go rates, as of 30 Sep 2025.
- Which is better, GPT-5.4 Mini or GLM 4.6?
- On published benchmarks, GPT-5.4 Mini. It leads on 10 of the 11 figures both models report, including coding ability (Coding Index), where it scores 56.1 against 45.8.
- Does GPT-5.4 Mini or GLM 4.6 offer batch pricing?
- GPT-5.4 Mini only. Its batch input costs $0.375, 50% off its standard rate, and GLM 4.6 lists no batch tier at all.
- Does GPT-5.4 Mini or GLM 4.6 have a bigger context window?
- GPT-5.4 Mini, at 400,000 tokens against 204,800. That is 195,200 more input tokens in a single request.