Cheaper overall
GLM 5.3 Flash
86% lower blended rate
Higher benchmark scores
GLM 5.3 Flash
leads on 9 of 11 figures
Bigger context window
GLM 5.3 Flash
1,000,000 vs 400,000 tokens
Cheaper cached input
GLM 5.3 Flash
$0.03 vs $0.075, 60% less
Where each one wins
GLM 5.3 Flash
- Cheaper input tokens $0.15 vs $0.75, 80% less
- Cheaper output tokens $0.50 vs $4.50, 89% less
- Cheaper cached input $0.03 vs $0.075, 60% less
- Cheaper on all four workloads
- Larger context window 1,000,000 tokens
- Higher overall intelligence (Intelligence Index) 41.9 vs 24.1
- Higher coding ability (Coding Index) 71.5 vs 56.1
- Higher multi-step tool use (Agentic Index) 51.2 vs 17.9
- Ahead on 6 more benchmark figures
GPT-5.4 Mini
- Higher scientific coding (SciCode) 52.1 vs 51.6
- Higher factual accuracy (Omniscience Accuracy) 37.5 vs 27.5
- Offers Batch and Flex pricing not listed for GLM 5.3 Flash
What four workloads cost
One run, and the same run a thousand times.
GLM 5.3 Flash is cheaper on all four workloads, by much the same margin each time โ about 5.7 times, from $0.0004 against $0.0030 on a chat turn to $0.0615 against $0.3135 on a long document.
| Workload | GLM 5.3 Flash | GPT-5.4 Mini | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0004 | $0.0030 | $0.4000 / $3.00 | GLM 5.3 Flash is 87% cheaper |
| RAG answer 10,000 in / 800 out | $0.0019 | $0.0111 | $1.90 / $11.10 | GLM 5.3 Flash is 83% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.0310 | $0.1380 | $31.00 / $138.00 | GLM 5.3 Flash is 78% cheaper |
| Long document 400,000 in / 3,000 out | $0.0615 | $0.3135 | $61.50 / $313.50 | GLM 5.3 Flash is 80% cheaper |
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | GLM 5.3 Flash | GPT-5.4 Mini | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.15 | $0.75 | GLM 5.3 Flash is 80% cheaper |
| Any prompt length | Cached input | $0.03 | $0.075 | GLM 5.3 Flash is 60% cheaper |
| Any prompt length | Output | $0.50 | $4.50 | GLM 5.3 Flash is 89% cheaper |
Blended: GLM 5.3 Flash $0.2375, GPT-5.4 Mini $1.6875 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
GLM 5.3 Flash priced by Z.ai, GPT-5.4 Mini by OpenAI. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.4 Mini lists Batch and Flex, at up to 50% off its own standard rate; GLM 5.3 Flash offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GLM 5.3 Flash | GPT-5.4 Mini | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.4 Mini | Not offered | $0.375 in / $2.25 out 50% off Standard | — |
| Flex only on GPT-5.4 Mini | Not offered | $0.375 in / $2.25 out 50% off Standard | — |
Benchmarks
They share 11 figures: GLM 5.3 Flash leads on 9, GPT-5.4 Mini on 2. The widest gap is 62.2 points, on answer reliability (Omniscience Non-Hallucination).
Overall intelligence
Intelligence Index โ Overall intelligence across reasoning, knowledge & math evals
GLM 5.3 Flash ahead by 17.8
Coding ability
Coding Index โ Coding ability across software-engineering evals
GLM 5.3 Flash ahead by 15.4
Multi-step tool use
Agentic Index โ Tool use & multi-step agent task performance
GLM 5.3 Flash ahead by 33.3
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GLM 5.3 Flash ahead by 3.7
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GLM 5.3 Flash ahead by 11.8
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GLM 5.3 Flash ahead by 3
Scientific coding
SciCode โ Research-level scientific coding tasks
GPT-5.4 Mini ahead by 0.5
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GLM 5.3 Flash ahead by 5.4
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
GLM 5.3 Flash ahead by 32.7
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GPT-5.4 Mini ahead by 10
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GLM 5.3 Flash ahead by 62.2
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GLM 5.3 Flash holds 600,000 more input tokens in one request. GLM 5.3 Flash also takes video.
| Specification | GLM 5.3 Flash | GPT-5.4 Mini |
|---|---|---|
| Context window | 1,000,000 tokens | 400,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Takes in | Text, image, video, file | Text, image, file |
| Puts out | Text | Text |
| Knowledge cutoff | โ | 31 Aug 2025 |
| Released | 26 Aug 2026 | 17 Mar 2026 |
| Status | Active | Active |
| Sold by | Perplexity, Z.ai | OpenAI, Perplexity |
FAQs
- Is GLM 5.3 Flash cheaper than GPT-5.4 Mini?
- Yes. A 1,000-token prompt with a 500-token reply costs $0.0004 on GLM 5.3 Flash and $0.0030 on GPT-5.4 Mini, and the same model is cheaper on every workload on this page.
- How much do GLM 5.3 Flash and GPT-5.4 Mini cost per 1M tokens?
- GLM 5.3 Flash costs $0.15 for input and $0.50 for output. GPT-5.4 Mini costs $0.75 and $4.50. Standard pay-as-you-go rates, as of 26 Aug 2026.
- Which is better, GLM 5.3 Flash or GPT-5.4 Mini?
- On published benchmarks, GLM 5.3 Flash. It leads on 9 of the 11 figures both models report, including overall intelligence (Intelligence Index), where it scores 41.9 against 24.1.
- Does GLM 5.3 Flash or GPT-5.4 Mini offer batch pricing?
- GPT-5.4 Mini only. Its batch input costs $0.375, 50% off its standard rate, and GLM 5.3 Flash lists no batch tier at all.
- Does GLM 5.3 Flash or GPT-5.4 Mini have a bigger context window?
- GLM 5.3 Flash, at 1,000,000 tokens against 400,000. That is 600,000 more input tokens in a single request.