Cheaper overall
GLM 4.7 FlashX
85% lower blended rate
Higher benchmark scores
GLM 4.7
leads on 14 of 16 figures
Bigger context window
GLM 4.7
204,800 vs 202,752 tokens
Cheaper cached input
GLM 4.7 FlashX
$0.01 vs $0.11, 91% less
Where each one wins
GLM 4.7 FlashX
- Cheaper input tokens $0.07 vs $0.60, 88% less
- Cheaper output tokens $0.40 vs $2.20, 82% less
- Cheaper cached input $0.01 vs $0.11, 91% less
- Cheaper on all four workloads
- Higher support-agent performance (ฯยฒ-Bench Telecom) 98.8 vs 95.9
- Higher Uicomponent Elo 1,216 vs 1,207
GLM 4.7
- Larger context window 204,800 tokens
- Longer maximum output 131,072 tokens
- Higher graduate-level science (GPQA Diamond) 85.9 vs 58.1
- Higher expert-exam performance (Humanity's Last Exam) 27.4 vs 7.6
- Higher instruction following (IFBench) 67.9 vs 60.8
- Ahead on 11 more benchmark figures
What four workloads cost
One run, and the same run a thousand times.
GLM 4.7 FlashX is cheaper on all four workloads, by much the same margin each time โ about 7.4 times, from $0.0003 against $0.0017 on a chat turn to $0.0144 against $0.1208 on a long document.
| Workload | GLM 4.7 FlashX | GLM 4.7 | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0003 | $0.0017 | $0.2700 / $1.70 | GLM 4.7 FlashX is 84% cheaper |
| RAG answer 10,000 in / 800 out | $0.0010 | $0.0078 | $1.02 / $7.76 | GLM 4.7 FlashX is 87% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.0144 | $0.1208 | $14.40 / $120.80 | GLM 4.7 FlashX is 88% cheaper |
| Long document 400,000 in / 3,000 out | GLM 4.7 FlashX's context window holds 202,752 tokens, so a 400,000-token prompt does not fit. | |||
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | GLM 4.7 FlashX | GLM 4.7 | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.07 | $0.60 | GLM 4.7 FlashX is 88% cheaper |
| Any prompt length | Cached input | $0.01 | $0.11 | GLM 4.7 FlashX is 91% cheaper |
| Any prompt length | Output | $0.40 | $2.20 | GLM 4.7 FlashX is 82% cheaper |
Blended: GLM 4.7 FlashX $0.1525, GLM 4.7 $1.00 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Both priced by Z.ai. Standard tier, pay-as-you-go.
Benchmarks
They share 16 figures: GLM 4.7 FlashX leads on 2, GLM 4.7 on 14. The widest gap is 105 points, on Svg Elo.
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GLM 4.7 ahead by 27.8
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GLM 4.7 ahead by 19.8
Instruction following
IFBench โ Precise following of detailed instructions
GLM 4.7 ahead by 7.1
Support-agent performance
ฯยฒ-Bench Telecom โ Tool-using agent tasks in a telecom support setting
GLM 4.7 FlashX ahead by 2.9
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GLM 4.7 ahead by 29.3
Command-line work
Terminal-Bench Hard โ Complex command-line and terminal workflows
GLM 4.7 ahead by 9.8
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GLM 4.7 ahead by 1.4
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GLM 4.7 ahead by 13.1
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GLM 4.7 ahead by 1
3d Elo
Design Arena head-to-head rating in the 3d category
GLM 4.7 ahead by 70
Codecategories Elo
Design Arena head-to-head rating in the codecategories category
GLM 4.7 ahead by 39
Dataviz Elo
Design Arena head-to-head rating in the dataviz category
GLM 4.7 ahead by 72
Gamedev Elo
Design Arena head-to-head rating in the gamedev category
GLM 4.7 ahead by 53
Svg Elo
Design Arena head-to-head rating in the svg category
GLM 4.7 ahead by 105
Uicomponent Elo
Design Arena head-to-head rating in the uicomponent category
GLM 4.7 FlashX ahead by 9
Website Elo
Design Arena head-to-head rating in the website category
GLM 4.7 ahead by 32
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GLM 4.7 holds 2,048 more input tokens in one request. Both list the same capabilities, so nothing here separates them.
| Specification | GLM 4.7 FlashX | GLM 4.7 |
|---|---|---|
| Context window | 202,752 tokens | 204,800 tokens |
| Max output | 128,000 tokens | 131,072 tokens |
| Takes in | Text | Text |
| Puts out | Text | Text |
| Released | 19 Jan 2026 | 22 Dec 2025 |
| Status | Active | Active |
| Sold by | Z.ai | Z.ai |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Yes |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Not listed | Not listed |
| Computer use | Not listed | Not listed |
| Open weights | Yes | Yes |
| Reasoning | Yes | Yes |
FAQs
- Is GLM 4.7 FlashX cheaper than GLM 4.7?
- Yes. A 1,000-token prompt with a 500-token reply costs $0.0003 on GLM 4.7 FlashX and $0.0017 on GLM 4.7, and the same model is cheaper on every workload on this page.
- How much do GLM 4.7 FlashX and GLM 4.7 cost per 1M tokens?
- GLM 4.7 FlashX costs $0.07 for input and $0.40 for output. GLM 4.7 costs $0.60 and $2.20. Standard pay-as-you-go rates, as of 19 Jan 2026.
- Which is better, GLM 4.7 FlashX or GLM 4.7?
- On published benchmarks, GLM 4.7. It leads on 14 of the 16 figures both models report, including graduate-level science (GPQA Diamond), where it scores 85.9 against 58.1.
- Does GLM 4.7 FlashX or GLM 4.7 have a bigger context window?
- GLM 4.7, at 204,800 tokens against 202,752. That is 2,048 more input tokens in a single request.