GLM 5.3 Flash vs GLM 5.3

GLM 5.3 Flash costs a fraction of what GLM 5.3 does: $0.15 against $1.40 per 1M for input, $0.50 against $4.40 for output, around 8.8 times as much. A single chat turn costs a fraction of a cent on both, so this only becomes real money at volume: a thousand long documents costs $61.50 against $573.20.

What the premium buys is 2.9 points of overall intelligence (Intelligence Index): 44.8 (max) against 41.9. Unless you genuinely need that lead, take GLM 5.3 Flash โ€” the premium here is steep.

Standard pay-as-you-go prices, as of 26 Aug 2026. Last updated 27 Sep 2026.

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Cheaper overall

GLM 5.3 Flash

89% lower blended rate

Higher benchmark scores

GLM 5.3

leads on 14 of 19 figures

Cheaper cached input

GLM 5.3 Flash

$0.03 vs $0.26, 88% less

Where each one wins

GLM 5.3 Flash

  • Cheaper input tokens $0.15 vs $1.40, 89% less
  • Cheaper output tokens $0.50 vs $4.40, 89% less
  • Cheaper cached input $0.03 vs $0.26, 88% less
  • Cheaper on all four workloads
  • Higher long-context reasoning (AA-LCR) 80 vs 79.7
  • Higher real-world task performance (GDPval) 57.7 vs 57.2
  • Higher answer reliability (Omniscience Non-Hallucination) 72.4 vs 70.4
  • Ahead on 2 more benchmark figures

GLM 5.3

  • Higher overall intelligence (Intelligence Index) 44.8 vs 41.9
  • Higher coding ability (Coding Index) 74.8 vs 71.5
  • Higher multi-step tool use (Agentic Index) 53.1 vs 51.2
  • Ahead on 11 more benchmark figures

What four workloads cost

One run, and the same run a thousand times.

GLM 5.3 Flash is cheaper on all four workloads, by much the same margin each time โ€” about 9.1 times, from $0.0004 against $0.0036 on a chat turn to $0.0615 against $0.5732 on a long document.

Workload GLM 5.3 Flash GLM 5.3 ร—1,000 Difference
Chat turn 1,000 in / 500 out $0.0004 $0.0036 $0.4000 / $3.60 GLM 5.3 Flash is 89% cheaper
RAG answer 10,000 in / 800 out $0.0019 $0.0175 $1.90 / $17.52 GLM 5.3 Flash is 89% cheaper
Agent loop 32,000 in / 700 out ร— 20 calls $0.0310 $0.2736 $31.00 / $273.60 GLM 5.3 Flash is 89% cheaper
Long document 400,000 in / 3,000 out $0.0615 $0.5732 $61.50 / $573.20 GLM 5.3 Flash is 89% cheaper

Price your own token counts for these two โ†’

Price per 1M tokens

Neither model changes its rate with prompt length, so these rates apply to every request.

Context band Price GLM 5.3 Flash GLM 5.3 Difference
Any prompt length Input $0.15 $1.40 GLM 5.3 Flash is 89% cheaper
Any prompt length Cached input $0.03 $0.26 GLM 5.3 Flash is 88% cheaper
Any prompt length Output $0.50 $4.40 GLM 5.3 Flash is 89% cheaper

Blended: GLM 5.3 Flash $0.2375, GLM 5.3 $2.15 per 1M โ€” one rate at 3:1 input to output, for comparing two models at a glance.

Both priced by Z.ai. Standard tier, pay-as-you-go.

Benchmarks

They share 19 figures: GLM 5.3 Flash leads on 5, GLM 5.3 on 14. The widest gap is 60 points, on Gamedev Elo.

Overall intelligence

Intelligence Index โ€” Overall intelligence across reasoning, knowledge & math evals

GLM 5.3 Flash 41.9
GLM 5.3 44.8

GLM 5.3 ahead by 2.9

Coding ability

Coding Index โ€” Coding ability across software-engineering evals

GLM 5.3 Flash 71.5
GLM 5.3 74.8

GLM 5.3 ahead by 3.3

Multi-step tool use

Agentic Index โ€” Tool use & multi-step agent task performance

GLM 5.3 Flash 51.2
GLM 5.3 53.1

GLM 5.3 ahead by 1.9

Graduate-level science

GPQA Diamond โ€” Graduate-level questions in biology, chemistry & physics

GLM 5.3 Flash 91.2
GLM 5.3 91.7

GLM 5.3 ahead by 0.5

Expert-exam performance

Humanity's Last Exam โ€” Expert-level questions across many academic domains

GLM 5.3 Flash 39.9
GLM 5.3 42.3

GLM 5.3 ahead by 2.4

Long-context reasoning

AA-LCR โ€” Long-context reasoning across large inputs

GLM 5.3 Flash 80
GLM 5.3 79.7

GLM 5.3 Flash ahead by 0.3

Scientific coding

SciCode โ€” Research-level scientific coding tasks

GLM 5.3 Flash 51.6
GLM 5.3 59

GLM 5.3 ahead by 7.4

Physics research reasoning

CritPt โ€” Unpublished physics research reasoning problems

GLM 5.3 Flash 15.4
GLM 5.3 19.1

GLM 5.3 ahead by 3.7

Real-world task performance

GDPval โ€” Economically valuable, real-world knowledge work

GLM 5.3 Flash 57.7
GLM 5.3 57.2

GLM 5.3 Flash ahead by 0.5

Factual accuracy

Omniscience Accuracy โ€” Breadth of factual knowledge across domains

GLM 5.3 Flash 27.5
GLM 5.3 33.9

GLM 5.3 ahead by 6.4

Answer reliability

Omniscience Non-Hallucination โ€” How reliably the model avoids fabricated answers

GLM 5.3 Flash 72.4
GLM 5.3 70.4

GLM 5.3 Flash ahead by 2

3d Elo

Design Arena head-to-head rating in the 3d category

GLM 5.3 Flash 1,337
GLM 5.3 1,375

GLM 5.3 ahead by 38

Asciiart Elo

Design Arena head-to-head rating in the asciiart category

GLM 5.3 Flash 1,272
GLM 5.3 1,218

GLM 5.3 Flash ahead by 54

Codecategories Elo

Design Arena head-to-head rating in the codecategories category

GLM 5.3 Flash 1,291
GLM 5.3 1,321

GLM 5.3 ahead by 30

Dataviz Elo

Design Arena head-to-head rating in the dataviz category

GLM 5.3 Flash 1,274
GLM 5.3 1,272

GLM 5.3 Flash ahead by 2

Gamedev Elo

Design Arena head-to-head rating in the gamedev category

GLM 5.3 Flash 1,295
GLM 5.3 1,355

GLM 5.3 ahead by 60

Svg Elo

Design Arena head-to-head rating in the svg category

GLM 5.3 Flash 1,299
GLM 5.3 1,307

GLM 5.3 ahead by 8

Uicomponent Elo

Design Arena head-to-head rating in the uicomponent category

GLM 5.3 Flash 1,328
GLM 5.3 1,331

GLM 5.3 ahead by 3

Website Elo

Design Arena head-to-head rating in the website category

GLM 5.3 Flash 1,280
GLM 5.3 1,309

GLM 5.3 ahead by 29

Each figure is that model's best published run, with the effort level named beside it โ€” where these scores come from.

Specifications

GLM 5.3 Flash also takes file, image and video. Only GLM 5.3 Flash lists computer use.

Specification GLM 5.3 Flash GLM 5.3
Context window 1,000,000 tokens 1,000,000 tokens
Max output 128,000 tokens 128,000 tokens
Takes in Text, image, video, file Text
Puts out Text Text
Released 26 Aug 2026 14 Aug 2026
Status Active Active
Sold by Perplexity, Z.ai Perplexity, Z.ai
Tool use Yes Yes
Structured outputs Yes Yes
Web search Not listed Not listed
Prompt caching Yes Yes
Code execution Yes Yes
Computer use Yes Not listed
Open weights Yes Yes
Reasoning Yes Yes

FAQs

Is GLM 5.3 Flash cheaper than GLM 5.3?
Yes. A 1,000-token prompt with a 500-token reply costs $0.0004 on GLM 5.3 Flash and $0.0036 on GLM 5.3, and the same model is cheaper on every workload on this page.
How much do GLM 5.3 Flash and GLM 5.3 cost per 1M tokens?
GLM 5.3 Flash costs $0.15 for input and $0.50 for output. GLM 5.3 costs $1.40 and $4.40. Standard pay-as-you-go rates, as of 26 Aug 2026.
Which is better, GLM 5.3 Flash or GLM 5.3?
On published benchmarks, GLM 5.3. It leads on 14 of the 19 figures both models report, including overall intelligence (Intelligence Index), where it scores 44.8 against 41.9.

Explore these models

GLM 5.3 Flash

Z.ai ยท 1,000,000-token context

Pricing · Calculator · Specifications

GLM 5.3

Z.ai ยท 1,000,000-token context

Pricing · Calculator · Specifications

Other comparisons