GLM 5.3 FlashX Pricing - Cost Calculator

GLM-5.3-FlashX is Z.aiโ€™s high-speed multimodal model for coding, visual understanding, tool use, long-context agents, and latency-sensitive production workflows.

GLM 5.3 FlashX, released 18 September 2026, is a multimodal model from Z.ai. It pairs a 1,048,576-token context window with up to 131,072 tokens of output, taking text and image input and returning text. Step-by-step reasoning is supported. It costs $0.37 per million input tokens through Z.ai.

Last updated Sep 18, 2026

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Model Details

Released
Sep 18, 2026
Context Length
1,048,576
Max Output
131,072
Modalities
Text Image Text
Capabilities
Tool use Structured outputs Prompt caching Code execution Open weights

Where is GLM 5.3 FlashX a perfect fit?

GLM-5.3-FlashX is Z.aiโ€™s high-speed serving variant of GLM-5.3-Flash, delivering the same multimodal intelligence with substantially faster inference for coding, agents, visual workflows, and production applications. Z.ai describes it as an inference/serving optimization of GLM-5.3-Flash, with reported peak speeds of up to 200 tokens/second.
- Low-latency coding assistants
- High-throughput autonomous agents
- Visual coding and screenshot understanding
- Browser and computer-use workflows
- Long-context document processing
- Tool-calling applications
- Real-time research and information processing
- Production applications where inference latency is important
GLM-5.3-FlashX is best treated as the high-speed serving variant of GLM-5.3-Flash, not as a separate foundation model. Its primary differentiator is latency and throughput while retaining the 320B/18B architecture, 1M-token context and multimodal capabilities of Flash.

Quick Model Estimate

(USD 0.3700 per 1M tokens)
(USD 1.2500 per 1M tokens)

Your GLM 5.3 FlashX Cost Estimate

๐Ÿ’ฐ Total Cost

โ€”

for 1000 input + 1000 output tokens

๐Ÿ“ฅ Input (1000 ร— $0.370000) โ€”
๐Ÿ“ค Output (1000 ร— $1.250000) โ€”

Cost Breakdown

๐Ÿ“ฅ Input ๐Ÿ“ค Output

Prices updated daily from official provider data.

Pricing

Provider โ†•
Modality โ†•
Service Tier โ†•
Input Price
(per 1M tokens)
โ†•
Output Price
(per 1M tokens)
โ†•
Cached Input
(per 1M tokens)
โ†•
Context Size โ†•
View
Z.ai LogoZ.aiTextStandard$0.3700$1.2500$0.07501,048,576 tokensโ†’

FAQs about GLM 5.3 FlashX

How much does GLM 5.3 FlashX cost per 1M tokens?

GLM 5.3 FlashX costs $0.37 per million input tokens, and $1.25 per million output tokens.

What does a typical workload cost with GLM 5.3 FlashX?

1,000 requests of 2,000 input and 500 output tokens each โ€” 2,000,000 input and 500,000 output tokens in total โ€” costs $1.365 with GLM 5.3 FlashX at its lowest rates. Use the calculator on this page for your own volumes.

Where can I use GLM 5.3 FlashX?

GLM 5.3 FlashX is available through Z.ai, from $0.37 per million input tokens.

What is GLM 5.3 FlashX's context window?

GLM 5.3 FlashX has a context window of 1,048,576 tokens, and returns up to 131,072 tokens in a single response.

What input types does GLM 5.3 FlashX support?

GLM 5.3 FlashX accepts text and image input, and returns text.