Gemini 3.8 Flash vs GLM 5.3 Flash

GLM 5.3 Flash costs a fraction of what Gemini 3.8 Flash does: $0.15 against $0.75 per 1M for input, $0.50 against $3.75 for output, around 5 times as much. Caching narrows the gap: over a thousand agent loops, where most of each prompt comes from cache, that is $31.00 against $127.50.

GLM 5.3 Flash also scores higher on overall intelligence (Intelligence Index), 41.9 against 40.9 (high), though only just, which makes price the stronger reason to prefer it.

Standard pay-as-you-go prices, as of 10 Sep 2026. Last updated 27 Sep 2026.

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Cheaper overall

GLM 5.3 Flash

84% lower blended rate

Higher benchmark scores

Gemini 3.8 Flash

leads on 10 of 16 figures

Bigger context window

Gemini 3.8 Flash

1,048,576 vs 1,000,000 tokens

Cheaper cached input

GLM 5.3 Flash

$0.03 vs $0.075, 60% less

Where each one wins

Gemini 3.8 Flash

  • Larger context window 1,048,576 tokens
  • Higher coding ability (Coding Index) 76.3 vs 71.5
  • Higher graduate-level science (GPQA Diamond) 95.3 vs 91.2
  • Higher expert-exam performance (Humanity's Last Exam) 47.8 vs 39.9
  • Ahead on 7 more benchmark figures

GLM 5.3 Flash

  • Cheaper input tokens $0.15 vs $0.75, 80% less
  • Cheaper output tokens $0.50 vs $3.75, 87% less
  • Cheaper cached input $0.03 vs $0.075, 60% less
  • Cheaper on all four workloads
  • Longer maximum output 128,000 tokens
  • Higher overall intelligence (Intelligence Index) 41.9 vs 40.9
  • Higher multi-step tool use (Agentic Index) 51.2 vs 40.2
  • Higher real-world task performance (GDPval) 57.7 vs 45.6
  • Ahead on 3 more benchmark figures

What four workloads cost

One run, and the same run a thousand times.

GLM 5.3 Flash is cheaper on all four workloads, by much the same margin each time โ€” about 5.3 times, from $0.0004 against $0.0026 on a chat turn to $0.0615 against $0.3113 on a long document.

Workload Gemini 3.8 Flash GLM 5.3 Flash ร—1,000 Difference
Chat turn 1,000 in / 500 out $0.0026 $0.0004 $2.63 / $0.4000 GLM 5.3 Flash is 85% cheaper
RAG answer 10,000 in / 800 out $0.0105 $0.0019 $10.50 / $1.90 GLM 5.3 Flash is 82% cheaper
Agent loop 32,000 in / 700 out ร— 20 calls $0.1275 $0.0310 $127.50 / $31.00 GLM 5.3 Flash is 76% cheaper
Long document 400,000 in / 3,000 out $0.3113 $0.0615 $311.25 / $61.50 GLM 5.3 Flash is 80% cheaper

Price your own token counts for these two โ†’

Price per 1M tokens

Neither model changes its rate with prompt length, so these rates apply to every request.

Context band Price Gemini 3.8 Flash GLM 5.3 Flash Difference
Any prompt length Input $0.75 $0.15 GLM 5.3 Flash is 80% cheaper
Any prompt length Cached input $0.075 $0.03 GLM 5.3 Flash is 60% cheaper
Any prompt length Output $3.75 $0.50 GLM 5.3 Flash is 87% cheaper

Blended: Gemini 3.8 Flash $1.50, GLM 5.3 Flash $0.2375 per 1M โ€” one rate at 3:1 input to output, for comparing two models at a glance.

Gemini 3.8 Flash priced by Perplexity, GLM 5.3 Flash by Z.ai. Standard tier, pay-as-you-go.

Benchmarks

They share 17 figures: Gemini 3.8 Flash leads on 10, GLM 5.3 Flash on 6. The widest gap is 31 points, on Gamedev Elo.

Overall intelligence

Intelligence Index โ€” Overall intelligence across reasoning, knowledge & math evals

Gemini 3.8 Flash 40.9
GLM 5.3 Flash 41.9

GLM 5.3 Flash ahead by 1

Coding ability

Coding Index โ€” Coding ability across software-engineering evals

Gemini 3.8 Flash 76.3
GLM 5.3 Flash 71.5

Gemini 3.8 Flash ahead by 4.8

Multi-step tool use

Agentic Index โ€” Tool use & multi-step agent task performance

Gemini 3.8 Flash 40.2
GLM 5.3 Flash 51.2

GLM 5.3 Flash ahead by 11

Graduate-level science

GPQA Diamond โ€” Graduate-level questions in biology, chemistry & physics

Gemini 3.8 Flash 95.3
GLM 5.3 Flash 91.2

Gemini 3.8 Flash ahead by 4.1

Expert-exam performance

Humanity's Last Exam โ€” Expert-level questions across many academic domains

Gemini 3.8 Flash 47.8
GLM 5.3 Flash 39.9

Gemini 3.8 Flash ahead by 7.9

Long-context reasoning

AA-LCR โ€” Long-context reasoning across large inputs

Gemini 3.8 Flash 81.3
GLM 5.3 Flash 80

Gemini 3.8 Flash ahead by 1.3

Scientific coding

SciCode โ€” Research-level scientific coding tasks

Gemini 3.8 Flash 56.6
GLM 5.3 Flash 51.6

Gemini 3.8 Flash ahead by 5

Physics research reasoning

CritPt โ€” Unpublished physics research reasoning problems

Gemini 3.8 Flash 18.3
GLM 5.3 Flash 15.4

Gemini 3.8 Flash ahead by 2.9

Real-world task performance

GDPval โ€” Economically valuable, real-world knowledge work

Gemini 3.8 Flash 45.6
GLM 5.3 Flash 57.7

GLM 5.3 Flash ahead by 12.1

Factual accuracy

Omniscience Accuracy โ€” Breadth of factual knowledge across domains

Gemini 3.8 Flash 54.6
GLM 5.3 Flash 27.5

Gemini 3.8 Flash ahead by 27.1

Answer reliability

Omniscience Non-Hallucination โ€” How reliably the model avoids fabricated answers

Gemini 3.8 Flash 44.8
GLM 5.3 Flash 72.4

GLM 5.3 Flash ahead by 27.6

3d Elo

Design Arena head-to-head rating in the 3d category

Gemini 3.8 Flash 1,316
GLM 5.3 Flash 1,337

GLM 5.3 Flash ahead by 21

Codecategories Elo

Design Arena head-to-head rating in the codecategories category

Gemini 3.8 Flash 1,315
GLM 5.3 Flash 1,291

Gemini 3.8 Flash ahead by 24

Dataviz Elo

Design Arena head-to-head rating in the dataviz category

Gemini 3.8 Flash 1,272
GLM 5.3 Flash 1,274

GLM 5.3 Flash ahead by 2

Gamedev Elo

Design Arena head-to-head rating in the gamedev category

Gemini 3.8 Flash 1,326
GLM 5.3 Flash 1,295

Gemini 3.8 Flash ahead by 31

Uicomponent Elo

Design Arena head-to-head rating in the uicomponent category

Gemini 3.8 Flash 1,328
GLM 5.3 Flash 1,328

Level

Website Elo

Design Arena head-to-head rating in the website category

Gemini 3.8 Flash 1,307
GLM 5.3 Flash 1,280

Gemini 3.8 Flash ahead by 27

Each figure is that model's best published run, with the effort level named beside it โ€” where these scores come from.

Specifications

Gemini 3.8 Flash holds 48,576 more input tokens in one request. Gemini 3.8 Flash also takes audio. Only GLM 5.3 Flash lists open weights.

Specification Gemini 3.8 Flash GLM 5.3 Flash
Context window 1,048,576 tokens 1,000,000 tokens
Max output 65,536 tokens 128,000 tokens
Takes in Text, audio, image, video, file Text, image, video, file
Puts out Text Text
Released 2 Sep 2026 26 Aug 2026
Status Active Active
Sold by Perplexity Perplexity, Z.ai
Tool use Yes Yes
Structured outputs Yes Yes
Web search Not listed Not listed
Prompt caching Yes Yes
Code execution Yes Yes
Computer use Yes Yes
Open weights Not listed Yes
Reasoning Yes Yes

FAQs

Is Gemini 3.8 Flash cheaper than GLM 5.3 Flash?
No โ€” the other way round. A 1,000-token prompt with a 500-token reply costs $0.0026 on Gemini 3.8 Flash and $0.0004 on GLM 5.3 Flash, and the same model is cheaper on every workload on this page.
How much do Gemini 3.8 Flash and GLM 5.3 Flash cost per 1M tokens?
Gemini 3.8 Flash costs $0.75 for input and $3.75 for output. GLM 5.3 Flash costs $0.15 and $0.50. Standard pay-as-you-go rates, as of 10 Sep 2026.
Which is better, Gemini 3.8 Flash or GLM 5.3 Flash?
On published benchmarks, GLM 5.3 Flash. It leads on 10 of the 17 figures both models report, including overall intelligence (Intelligence Index), where it scores 41.9 against 40.9.
Does Gemini 3.8 Flash or GLM 5.3 Flash have a bigger context window?
Gemini 3.8 Flash, at 1,048,576 tokens against 1,000,000. That is 48,576 more input tokens in a single request.

Explore these models

Gemini 3.8 Flash

Perplexity ยท 1,048,576-token context

Pricing · Calculator · Specifications

GLM 5.3 Flash

Z.ai ยท 1,000,000-token context

Pricing · Calculator · Specifications

Other comparisons