Cheaper overall
GLM 5.3
28% lower blended rate
Bigger context window
GLM 5.3
1,000,000 vs 500,000 tokens
Cheaper cached input
GLM 5.3
$0.26 vs $0.50, 48% less
Where each one wins
Grok-4.7
- Higher overall intelligence (Intelligence Index) 46.4 vs 44.8
- Higher expert-exam performance (Humanity's Last Exam) 43.1 vs 42.3
- Higher real-world task performance (GDPval) 59.8 vs 57.2
- Ahead on 2 more benchmark figures
GLM 5.3
- Cheaper input tokens $1.40 vs $2.00, 30% less
- Cheaper output tokens $4.40 vs $6.00, 27% less
- Cheaper cached input $0.26 vs $0.50, 48% less
- Cheaper on all four workloads
- Larger context window 1,000,000 tokens
- Higher long-context reasoning (AA-LCR) 79.7 vs 76.7
- Higher scientific coding (SciCode) 59 vs 57.4
- Higher physics research reasoning (CritPt) 19.1 vs 17.7
- Ahead on 2 more benchmark figures
What four workloads cost
One run, and the same run a thousand times.
The gap opens on long prompts: GLM 5.3 costs $0.5732 against $1.64, a difference of $1,062.80 over a thousand runs.
| Workload | Grok-4.7 | GLM 5.3 | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0050 | $0.0036 | $5.00 / $3.60 | GLM 5.3 is 28% cheaper |
| RAG answer 10,000 in / 800 out | $0.0248 | $0.0175 | $24.80 / $17.52 | GLM 5.3 is 29% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.4640 | $0.2736 | $464.00 / $273.60 | GLM 5.3 is 41% cheaper |
| Long document 400,000 in / 3,000 out | $1.64 | $0.5732 | $1,636.00 / $573.20 | GLM 5.3 is 65% cheaper |
Price per 1M tokens
Above 199,999 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: Grok-4.7's input goes from $2.00 to $4.00. It does not change which of the two is cheaper.
| Context band | Price | Grok-4.7 | GLM 5.3 | Difference |
|---|---|---|---|---|
| Prompts up to 199,999 tokens | Input | $2.00 | $1.40 | GLM 5.3 is 30% cheaper |
| Prompts up to 199,999 tokens | Cached input | $0.50 | $0.26 | GLM 5.3 is 48% cheaper |
| Prompts up to 199,999 tokens | Output | $6.00 | $4.40 | GLM 5.3 is 27% cheaper |
| Prompts up to 200,000 tokens | Input | $4.00 | $1.40 | GLM 5.3 is 65% cheaper |
| Prompts up to 200,000 tokens | Cached input | $1.00 | $0.26 | GLM 5.3 is 74% cheaper |
| Prompts up to 200,000 tokens | Output | $12.00 | $4.40 | GLM 5.3 is 63% cheaper |
| Prompts over 200,000 tokens | Input | $4.00 | $1.40 | GLM 5.3 is 65% cheaper |
| Prompts over 200,000 tokens | Cached input | $1.00 | $0.26 | GLM 5.3 is 74% cheaper |
| Prompts over 200,000 tokens | Output | $12.00 | $4.40 | GLM 5.3 is 63% cheaper |
Blended: Grok-4.7 $3.00, GLM 5.3 $2.15 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Grok-4.7 priced by Perplexity, GLM 5.3 by Z.ai. Standard tier, pay-as-you-go.
Benchmarks
They share 10 figures: Grok-4.7 leads on 5, GLM 5.3 on 5. The widest gap is 85 points, on Website Elo.
Overall intelligence
Intelligence Index โ Overall intelligence across reasoning, knowledge & math evals
Grok-4.7 ahead by 1.6
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
Grok-4.7 ahead by 0.8
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GLM 5.3 ahead by 3
Scientific coding
SciCode โ Research-level scientific coding tasks
GLM 5.3 ahead by 1.6
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GLM 5.3 ahead by 1.4
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
Grok-4.7 ahead by 2.6
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
Grok-4.7 ahead by 13.5
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
Grok-4.7 ahead by 0.3
Gamedev Elo
Design Arena head-to-head rating in the gamedev category
GLM 5.3 ahead by 47
Website Elo
Design Arena head-to-head rating in the website category
GLM 5.3 ahead by 85
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GLM 5.3 holds 500,000 more input tokens in one request. Grok-4.7 also takes file and image. Only GLM 5.3 lists open weights.
| Specification | Grok-4.7 | GLM 5.3 |
|---|---|---|
| Context window | 500,000 tokens | 1,000,000 tokens |
| Max output | โ | 128,000 tokens |
| Takes in | Text, image, file | Text |
| Puts out | Text | Text |
| Knowledge cutoff | 31 May 2026 | โ |
| Released | 21 Sep 2026 | 14 Aug 2026 |
| Status | Active | Active |
| Sold by | Perplexity | Perplexity, Z.ai |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Yes |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Yes | Yes |
| Computer use | Not listed | Not listed |
| Open weights | Not listed | Yes |
| Reasoning | Yes | Yes |
FAQs
- Is Grok-4.7 cheaper than GLM 5.3?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0050 on Grok-4.7 and $0.0036 on GLM 5.3, and the same model is cheaper on every workload on this page.
- How much do Grok-4.7 and GLM 5.3 cost per 1M tokens?
- Grok-4.7 costs $2.00 for input and $6.00 for output. GLM 5.3 costs $1.40 and $4.40. Standard pay-as-you-go rates, as of 23 Sep 2026.
- Which is better, Grok-4.7 or GLM 5.3?
- On published benchmarks, Grok-4.7. It leads on 5 of the 10 figures both models report, including overall intelligence (Intelligence Index), where it scores 46.4 against 44.8.
- Does Grok-4.7 or GLM 5.3 have a bigger context window?
- GLM 5.3, at 1,000,000 tokens against 500,000. That is 500,000 more input tokens in a single request.