GLM 4.7 FlashX vs GLM 4.7

GLM 4.7 FlashX costs a fraction of what GLM 4.7 does: $0.07 against $0.60 per 1M for input, $0.40 against $2.20 for output, around 5.5 times as much. Caching widens the gap: over a thousand agent loops, where most of each prompt comes from cache, that is $14.40 against $120.80.

What the premium buys is 27.8 points of graduate-level science (GPQA Diamond): 85.9 (Reasoning) against 58.1 (GLM-4.7-Flash). Unless you genuinely need that lead, take GLM 4.7 FlashX โ€” the premium here is steep.

Standard pay-as-you-go prices, as of 19 Jan 2026. Last updated 29 Sep 2026.

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Cheaper overall

GLM 4.7 FlashX

85% lower blended rate

Higher benchmark scores

GLM 4.7

leads on 14 of 16 figures

Bigger context window

GLM 4.7

204,800 vs 202,752 tokens

Cheaper cached input

GLM 4.7 FlashX

$0.01 vs $0.11, 91% less

Where each one wins

GLM 4.7 FlashX

  • Cheaper input tokens $0.07 vs $0.60, 88% less
  • Cheaper output tokens $0.40 vs $2.20, 82% less
  • Cheaper cached input $0.01 vs $0.11, 91% less
  • Cheaper on all four workloads
  • Higher support-agent performance (ฯ„ยฒ-Bench Telecom) 98.8 vs 95.9
  • Higher Uicomponent Elo 1,216 vs 1,207

GLM 4.7

  • Larger context window 204,800 tokens
  • Longer maximum output 131,072 tokens
  • Higher graduate-level science (GPQA Diamond) 85.9 vs 58.1
  • Higher expert-exam performance (Humanity's Last Exam) 27.4 vs 7.6
  • Higher instruction following (IFBench) 67.9 vs 60.8
  • Ahead on 11 more benchmark figures

What four workloads cost

One run, and the same run a thousand times.

GLM 4.7 FlashX is cheaper on all four workloads, by much the same margin each time โ€” about 7.4 times, from $0.0003 against $0.0017 on a chat turn to $0.0144 against $0.1208 on a long document.

Workload GLM 4.7 FlashX GLM 4.7 ร—1,000 Difference
Chat turn 1,000 in / 500 out $0.0003 $0.0017 $0.2700 / $1.70 GLM 4.7 FlashX is 84% cheaper
RAG answer 10,000 in / 800 out $0.0010 $0.0078 $1.02 / $7.76 GLM 4.7 FlashX is 87% cheaper
Agent loop 32,000 in / 700 out ร— 20 calls $0.0144 $0.1208 $14.40 / $120.80 GLM 4.7 FlashX is 88% cheaper
Long document 400,000 in / 3,000 out GLM 4.7 FlashX's context window holds 202,752 tokens, so a 400,000-token prompt does not fit.

Price your own token counts for these two โ†’

Price per 1M tokens

Neither model changes its rate with prompt length, so these rates apply to every request.

Context band Price GLM 4.7 FlashX GLM 4.7 Difference
Any prompt length Input $0.07 $0.60 GLM 4.7 FlashX is 88% cheaper
Any prompt length Cached input $0.01 $0.11 GLM 4.7 FlashX is 91% cheaper
Any prompt length Output $0.40 $2.20 GLM 4.7 FlashX is 82% cheaper

Blended: GLM 4.7 FlashX $0.1525, GLM 4.7 $1.00 per 1M โ€” one rate at 3:1 input to output, for comparing two models at a glance.

Both priced by Z.ai. Standard tier, pay-as-you-go.

Benchmarks

They share 16 figures: GLM 4.7 FlashX leads on 2, GLM 4.7 on 14. The widest gap is 105 points, on Svg Elo.

Graduate-level science

GPQA Diamond โ€” Graduate-level questions in biology, chemistry & physics

GLM 4.7 FlashX 58.1
GLM 4.7 85.9

GLM 4.7 ahead by 27.8

Expert-exam performance

Humanity's Last Exam โ€” Expert-level questions across many academic domains

GLM 4.7 FlashX 7.6
GLM 4.7 27.4

GLM 4.7 ahead by 19.8

Instruction following

IFBench โ€” Precise following of detailed instructions

GLM 4.7 FlashX 60.8
GLM 4.7 67.9

GLM 4.7 ahead by 7.1

Support-agent performance

ฯ„ยฒ-Bench Telecom โ€” Tool-using agent tasks in a telecom support setting

GLM 4.7 FlashX 98.8
GLM 4.7 95.9

GLM 4.7 FlashX ahead by 2.9

Long-context reasoning

AA-LCR โ€” Long-context reasoning across large inputs

GLM 4.7 FlashX 41.7
GLM 4.7 71

GLM 4.7 ahead by 29.3

Command-line work

Terminal-Bench Hard โ€” Complex command-line and terminal workflows

GLM 4.7 FlashX 22
GLM 4.7 31.8

GLM 4.7 ahead by 9.8

Physics research reasoning

CritPt โ€” Unpublished physics research reasoning problems

GLM 4.7 FlashX 0.3
GLM 4.7 1.7

GLM 4.7 ahead by 1.4

Factual accuracy

Omniscience Accuracy โ€” Breadth of factual knowledge across domains

GLM 4.7 FlashX 16.2
GLM 4.7 29.3

GLM 4.7 ahead by 13.1

Answer reliability

Omniscience Non-Hallucination โ€” How reliably the model avoids fabricated answers

GLM 4.7 FlashX 6.1
GLM 4.7 7.1

GLM 4.7 ahead by 1

3d Elo

Design Arena head-to-head rating in the 3d category

GLM 4.7 FlashX 1,139
GLM 4.7 1,209

GLM 4.7 ahead by 70

Codecategories Elo

Design Arena head-to-head rating in the codecategories category

GLM 4.7 FlashX 1,187
GLM 4.7 1,226

GLM 4.7 ahead by 39

Dataviz Elo

Design Arena head-to-head rating in the dataviz category

GLM 4.7 FlashX 1,134
GLM 4.7 1,206

GLM 4.7 ahead by 72

Gamedev Elo

Design Arena head-to-head rating in the gamedev category

GLM 4.7 FlashX 1,149
GLM 4.7 1,202

GLM 4.7 ahead by 53

Svg Elo

Design Arena head-to-head rating in the svg category

GLM 4.7 FlashX 1,046
GLM 4.7 1,151

GLM 4.7 ahead by 105

Uicomponent Elo

Design Arena head-to-head rating in the uicomponent category

GLM 4.7 FlashX 1,216
GLM 4.7 1,207

GLM 4.7 FlashX ahead by 9

Website Elo

Design Arena head-to-head rating in the website category

GLM 4.7 FlashX 1,202
GLM 4.7 1,234

GLM 4.7 ahead by 32

Each figure is that model's best published run, with the effort level named beside it โ€” where these scores come from.

Specifications

GLM 4.7 holds 2,048 more input tokens in one request. Both list the same capabilities, so nothing here separates them.

Specification GLM 4.7 FlashX GLM 4.7
Context window 202,752 tokens 204,800 tokens
Max output 128,000 tokens 131,072 tokens
Takes in Text Text
Puts out Text Text
Released 19 Jan 2026 22 Dec 2025
Status Active Active
Sold by Z.ai Z.ai
Tool use Yes Yes
Structured outputs Yes Yes
Web search Not listed Not listed
Prompt caching Yes Yes
Code execution Not listed Not listed
Computer use Not listed Not listed
Open weights Yes Yes
Reasoning Yes Yes

FAQs

Is GLM 4.7 FlashX cheaper than GLM 4.7?
Yes. A 1,000-token prompt with a 500-token reply costs $0.0003 on GLM 4.7 FlashX and $0.0017 on GLM 4.7, and the same model is cheaper on every workload on this page.
How much do GLM 4.7 FlashX and GLM 4.7 cost per 1M tokens?
GLM 4.7 FlashX costs $0.07 for input and $0.40 for output. GLM 4.7 costs $0.60 and $2.20. Standard pay-as-you-go rates, as of 19 Jan 2026.
Which is better, GLM 4.7 FlashX or GLM 4.7?
On published benchmarks, GLM 4.7. It leads on 14 of the 16 figures both models report, including graduate-level science (GPQA Diamond), where it scores 85.9 against 58.1.
Does GLM 4.7 FlashX or GLM 4.7 have a bigger context window?
GLM 4.7, at 204,800 tokens against 202,752. That is 2,048 more input tokens in a single request.

Explore these models

GLM 4.7 FlashX

Z.ai ยท 202,752-token context

Pricing · Calculator · Specifications

GLM 4.7

Z.ai ยท 204,800-token context

Pricing · Calculator · Specifications

Other comparisons