GLM-5.3-Flash Pricing - Cost Calculator

GLM-5.3-Flash is Z.aiโ€™s efficient multimodal frontier model for coding, autonomous agents, visual workflows, document processing, research, and professional automation.

GLM-5.3-Flash is a multimodal model from Z.ai, released 26 August 2026. It pairs a 1,000,000-token context window with up to 128,000 tokens of output, accepting text, image, video and file and returning text. Step-by-step reasoning is supported. It costs $0.15 per million input tokens through Perplexity.

Last updated Sep 15, 2026

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Model Details

Released
Aug 26, 2026
Context Length
1,000,000
Max Output
128,000
Modalities
Text Image Video File Text
Capabilities
Tool use Structured outputs Prompt caching Code execution Computer use Open weights

Where is GLM-5.3-Flash a perfect fit?

GLM-5.3-Flash is Z.aiโ€™s highly efficient open-weight multimodal model, combining frontier coding and agentic intelligence with native visual understanding, long context, and exceptionally low inference costs.
- High-volume coding agents
- Autonomous software engineering
- Visual coding and UI development
- Browser and computer-use workflows
- Document and spreadsheet processing
- Financial research and analysis
- Multimodal research agents
- PPTX, PDF, DOCX and XLSX workflows
- Long-context analysis
- Cost-sensitive production agents

GLM-5.3-Flash uses 320B total parameters with only 18B activated, and its hybrid sparse/linear-attention architecture reduces attention computation by roughly 3ร— and KV-cache requirements by 4.4ร— versus GLM-5.3.
Z.ai reports a score of 57 on the Artificial Analysis Intelligence Index v4.1.1 at approximately $0.045 per task, positioning it strongly on the intelligence-to-cost frontier.

Quick Model Estimate

(USD 0.1500 per 1M tokens)
(USD 0.5000 per 1M tokens)

Your GLM-5.3-Flash Cost Estimate

๐Ÿ’ฐ Total Cost

โ€”

for 1000 input + 1000 output tokens

๐Ÿ“ฅ Input (1000 ร— $0.150000) โ€”
๐Ÿ“ค Output (1000 ร— $0.500000) โ€”

Cost Breakdown

๐Ÿ“ฅ Input ๐Ÿ“ค Output

Prices updated daily from official provider data.

Pricing

Provider โ†•
Modality โ†•
Service Tier โ†•
Input Price
(per 1M tokens)
โ†•
Output Price
(per 1M tokens)
โ†•
Cached Input
(per 1M tokens)
โ†•
Context Size โ†•
View
Perplexity LogoPerplexityTextStandard$0.1500$0.5000$0.03001,000,000 tokensโ†’

Benchmarks

Scores from standardized evaluations by Artificial Analysis and Design Arena. Higher is better โ€” the indices summarize overall ability, while the detailed scores break down performance on individual benchmarks.

Artificial Analysis

Higher is better ยท benchmarked by Artificial Analysis
Detailed scores
GLM-5.3-Flash
41.9
Intelligence Index
Overall intelligence across reasoning, knowledge & math evals
Better than 38% of 9 models
71.5
Coding Index
Coding ability across software-engineering evals
Better than 44% of 10 models
51.2
Agentic Index
Tool use & multi-step agent task performance
Better than 63% of 9 models
Reasoning
GPQA Diamond 91.2%

Graduate-level questions in biology, chemistry & physics

Humanity's Last Exam 39.9%

Expert-level questions across many academic domains

CritPt 15.4%

Unpublished physics research reasoning problems

AA-LCR 80%

Long-context reasoning across large inputs

Coding
SciCode 51.6%

Research-level scientific coding tasks

Knowledge
GDPval 57.7%

Economically valuable, real-world knowledge work

Omniscience Accuracy 27.5%

Breadth of factual knowledge across domains

Omniscience Non-Hallucination 72.4%

How reliably the model avoids fabricated answers

Design Arena

Elo rating by arena & category
Variant Arena Category Elo Win % Percentile Avg time (ms)
notebook models 3d 1,355 60% 94 449,565
notebook models asciiart 1,285 55.7% 87 408,625
notebook models codecategories 1,298 49.8% 88 361,371
notebook models dataviz 1,275 50.7% 84 338,814
notebook models gamedev 1,308 47.4% 90 578,721
notebook models svg 1,312 57.1% 93 188,916
notebook models uicomponent 1,342 56.6% 95 281,303
notebook models website 1,285 48.5% 86 313,229

FAQs about GLM-5.3-Flash

How much does GLM-5.3-Flash cost per 1M tokens?

GLM-5.3-Flash costs $0.15 per million input tokens, and $0.50 per million output tokens.

What does a typical workload cost with GLM-5.3-Flash?

1,000 requests of 2,000 input and 500 output tokens each โ€” 2,000,000 input and 500,000 output tokens in total โ€” costs $0.55 with GLM-5.3-Flash at its lowest rates. Use the calculator on this page for your own volumes.

Where can I use GLM-5.3-Flash?

GLM-5.3-Flash is available through Perplexity, from $0.15 per million input tokens.

What is GLM-5.3-Flash's context window?

GLM-5.3-Flash has a context window of 1,000,000 tokens, and returns up to 128,000 tokens in a single response.

What input types does GLM-5.3-Flash support?

GLM-5.3-Flash accepts text, image, video and file input, and returns text.