GLM 4.6V Flash Pricing - Cost Calculator

GLM-4.6V-Flash is Z.aiโ€™s lightweight open-source vision-language model for fast multimodal reasoning, visual understanding, document analysis, GUI agents, and tool-enabled applications.

GLM 4.6V Flash is a video-generation model from Z.ai, released 8 December 2025. It is free to use through Z.ai, with no charge for input tokens. Text, image, video and file input goes in; text comes out. It reasons step by step.

Last updated Oct 1, 2026

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Model Details

Released
Dec 8, 2025
Context Length
128,000
Max Output
32,000
Modalities
Text Image Video File → Text
Capabilities
Open weights

Where is GLM 4.6V Flash a perfect fit?

GLM-4.6V-Flash is a 9B multimodal model optimized for local deployment and low latency, supporting images, videos, documents, visual reasoning, native function calling, grounding, and GUI automation. It can be a perfect fit in below scenarios:
- High-volume image understanding
- OCR and document extraction
- Charts, tables and diagrams
- Screenshot and UI understanding
- GUI agents
- Visual grounding and object localization
- Video understanding
- Multimodal RAG
- Local/private vision applications
- Low-latency multimodal agents
- Visual tool calling

Quick Model Estimate

(Free)
(Free)

Your GLM 4.6V Flash Cost Estimate

๐Ÿ’ฐ Total Cost

โ€”

for 1000 input + 1000 output tokens

๐Ÿ“ฅ Input (1000 ร— Free) โ€”
๐Ÿ“ค Output (1000 ร— Free) โ€”

Cost Breakdown

๐Ÿ“ฅ Input ๐Ÿ“ค Output

Prices are watched for changes and checked against the provider's own page.

Pricing

Provider โ†•
Modality โ†•
Service Tier โ†•
Input Price
(per 1M tokens)
โ†•
Output Price
(per 1M tokens)
โ†•
Cached Input
(per 1M tokens)
โ†•
Context Size โ†•
View
Z.ai LogoZ.aiTextStandardFreeFreeFree128,000 tokensโ†’

FAQs about GLM 4.6V Flash

How much does GLM 4.6V Flash cost per 1M tokens?

GLM 4.6V Flash costs nothing for input tokens, and nothing for output tokens.

What does a typical workload cost with GLM 4.6V Flash?

1,000 requests of 2,000 input and 500 output tokens each โ€” 2,000,000 input and 500,000 output tokens in total โ€” costs nothing with GLM 4.6V Flash at its lowest rates. Use the calculator on this page for your own volumes.

Where can I use GLM 4.6V Flash?

GLM 4.6V Flash is available through Z.ai, free of charge for input tokens.

What is GLM 4.6V Flash's context window?

GLM 4.6V Flash has a context window of 128,000 tokens, and returns up to 32,000 tokens in a single response.

What input types does GLM 4.6V Flash support?

GLM 4.6V Flash accepts text, image, video and file input, and returns text.