Cheaper overall
GLM 5.3 Flash
84% lower blended rate
Higher benchmark scores
Gemini 3.8 Flash
leads on 10 of 16 figures
Bigger context window
Gemini 3.8 Flash
1,048,576 vs 1,000,000 tokens
Cheaper cached input
GLM 5.3 Flash
$0.03 vs $0.075, 60% less
Where each one wins
Gemini 3.8 Flash
- Larger context window 1,048,576 tokens
- Higher coding ability (Coding Index) 76.3 vs 71.5
- Higher graduate-level science (GPQA Diamond) 95.3 vs 91.2
- Higher expert-exam performance (Humanity's Last Exam) 47.8 vs 39.9
- Ahead on 7 more benchmark figures
GLM 5.3 Flash
- Cheaper input tokens $0.15 vs $0.75, 80% less
- Cheaper output tokens $0.50 vs $3.75, 87% less
- Cheaper cached input $0.03 vs $0.075, 60% less
- Cheaper on all four workloads
- Longer maximum output 128,000 tokens
- Higher overall intelligence (Intelligence Index) 41.9 vs 40.9
- Higher multi-step tool use (Agentic Index) 51.2 vs 40.2
- Higher real-world task performance (GDPval) 57.7 vs 45.6
- Ahead on 3 more benchmark figures
What four workloads cost
One run, and the same run a thousand times.
GLM 5.3 Flash is cheaper on all four workloads, by much the same margin each time โ about 5.3 times, from $0.0004 against $0.0026 on a chat turn to $0.0615 against $0.3113 on a long document.
| Workload | Gemini 3.8 Flash | GLM 5.3 Flash | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0026 | $0.0004 | $2.63 / $0.4000 | GLM 5.3 Flash is 85% cheaper |
| RAG answer 10,000 in / 800 out | $0.0105 | $0.0019 | $10.50 / $1.90 | GLM 5.3 Flash is 82% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1275 | $0.0310 | $127.50 / $31.00 | GLM 5.3 Flash is 76% cheaper |
| Long document 400,000 in / 3,000 out | $0.3113 | $0.0615 | $311.25 / $61.50 | GLM 5.3 Flash is 80% cheaper |
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | Gemini 3.8 Flash | GLM 5.3 Flash | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.75 | $0.15 | GLM 5.3 Flash is 80% cheaper |
| Any prompt length | Cached input | $0.075 | $0.03 | GLM 5.3 Flash is 60% cheaper |
| Any prompt length | Output | $3.75 | $0.50 | GLM 5.3 Flash is 87% cheaper |
Blended: Gemini 3.8 Flash $1.50, GLM 5.3 Flash $0.2375 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Gemini 3.8 Flash priced by Perplexity, GLM 5.3 Flash by Z.ai. Standard tier, pay-as-you-go.
Benchmarks
They share 17 figures: Gemini 3.8 Flash leads on 10, GLM 5.3 Flash on 6. The widest gap is 31 points, on Gamedev Elo.
Overall intelligence
Intelligence Index โ Overall intelligence across reasoning, knowledge & math evals
GLM 5.3 Flash ahead by 1
Coding ability
Coding Index โ Coding ability across software-engineering evals
Gemini 3.8 Flash ahead by 4.8
Multi-step tool use
Agentic Index โ Tool use & multi-step agent task performance
GLM 5.3 Flash ahead by 11
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
Gemini 3.8 Flash ahead by 4.1
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
Gemini 3.8 Flash ahead by 7.9
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
Gemini 3.8 Flash ahead by 1.3
Scientific coding
SciCode โ Research-level scientific coding tasks
Gemini 3.8 Flash ahead by 5
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
Gemini 3.8 Flash ahead by 2.9
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
GLM 5.3 Flash ahead by 12.1
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
Gemini 3.8 Flash ahead by 27.1
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GLM 5.3 Flash ahead by 27.6
3d Elo
Design Arena head-to-head rating in the 3d category
GLM 5.3 Flash ahead by 21
Codecategories Elo
Design Arena head-to-head rating in the codecategories category
Gemini 3.8 Flash ahead by 24
Dataviz Elo
Design Arena head-to-head rating in the dataviz category
GLM 5.3 Flash ahead by 2
Gamedev Elo
Design Arena head-to-head rating in the gamedev category
Gemini 3.8 Flash ahead by 31
Uicomponent Elo
Design Arena head-to-head rating in the uicomponent category
Level
Website Elo
Design Arena head-to-head rating in the website category
Gemini 3.8 Flash ahead by 27
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
Gemini 3.8 Flash holds 48,576 more input tokens in one request. Gemini 3.8 Flash also takes audio. Only GLM 5.3 Flash lists open weights.
| Specification | Gemini 3.8 Flash | GLM 5.3 Flash |
|---|---|---|
| Context window | 1,048,576 tokens | 1,000,000 tokens |
| Max output | 65,536 tokens | 128,000 tokens |
| Takes in | Text, audio, image, video, file | Text, image, video, file |
| Puts out | Text | Text |
| Released | 2 Sep 2026 | 26 Aug 2026 |
| Status | Active | Active |
| Sold by | Perplexity | Perplexity, Z.ai |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Yes |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Yes | Yes |
| Computer use | Yes | Yes |
| Open weights | Not listed | Yes |
| Reasoning | Yes | Yes |
FAQs
- Is Gemini 3.8 Flash cheaper than GLM 5.3 Flash?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0026 on Gemini 3.8 Flash and $0.0004 on GLM 5.3 Flash, and the same model is cheaper on every workload on this page.
- How much do Gemini 3.8 Flash and GLM 5.3 Flash cost per 1M tokens?
- Gemini 3.8 Flash costs $0.75 for input and $3.75 for output. GLM 5.3 Flash costs $0.15 and $0.50. Standard pay-as-you-go rates, as of 10 Sep 2026.
- Which is better, Gemini 3.8 Flash or GLM 5.3 Flash?
- On published benchmarks, GLM 5.3 Flash. It leads on 10 of the 17 figures both models report, including overall intelligence (Intelligence Index), where it scores 41.9 against 40.9.
- Does Gemini 3.8 Flash or GLM 5.3 Flash have a bigger context window?
- Gemini 3.8 Flash, at 1,048,576 tokens against 1,000,000. That is 48,576 more input tokens in a single request.