GLM 4.7 FlashX Pricing - Cost Calculator

GLM-4.7-FlashX is Z.aiโ€™s ultra-fast lightweight reasoning model for coding, agents, tool calling, and high-volume low-latency AI applications.

GLM 4.7 FlashX is a coding model from Z.ai, released 19 January 2026. It costs $0.07 per million input tokens through Z.ai. The context window holds 202,752 tokens, with up to 128,000 returned per response; the model takes text input and returns text. It reasons step by step.

Last updated Sep 29, 2026

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Model Details

Released
Jan 19, 2026
Context Length
202,752
Max Output
128,000
Modalities
Text → Text
Capabilities
Tool use Structured outputs Prompt caching Open weights

Where is GLM 4.7 FlashX a perfect fit?

GLM-4.7-FlashX combines a 30B-A3B MoE architecture with strong reasoning and coding capabilities, delivering substantially faster inference and low API costs for production workloads. It can be suitable for below use cases:
- Low-latency applications โ€” fast interactive AI responses
- Coding agents โ€” lightweight autonomous software-development agents
- Tool-calling workflows โ€” high-volume function and API calls
- Reasoning tasks โ€” mathematics, analysis, and structured problem solving
- Production chat โ€” inexpensive high-throughput conversational workloads
- Sub-agents โ€” economical worker models inside larger agent systems
- Self-hosted AI โ€” efficient deployment compared with large frontier models

Quick Model Estimate

(USD 0.0700 per 1M tokens)
(USD 0.4000 per 1M tokens)

Your GLM 4.7 FlashX Cost Estimate

๐Ÿ’ฐ Total Cost

โ€”

for 1000 input + 1000 output tokens

๐Ÿ“ฅ Input (1000 ร— $0.070000) โ€”
๐Ÿ“ค Output (1000 ร— $0.400000) โ€”

Cost Breakdown

๐Ÿ“ฅ Input ๐Ÿ“ค Output

Prices are watched for changes and checked against the provider's own page.

Pricing

Provider โ†•
Modality โ†•
Service Tier โ†•
Input Price
(per 1M tokens)
โ†•
Output Price
(per 1M tokens)
โ†•
Cached Input
(per 1M tokens)
โ†•
Context Size โ†•
View
Z.ai LogoZ.aiTextStandard$0.0700$0.4000$0.0100202,752 tokensโ†’

Benchmarks

Scores from standardized evaluations by Artificial Analysis and Design Arena. Higher is better โ€” the indices summarize overall ability, while the detailed scores break down performance on individual benchmarks.

Artificial Analysis

Higher is better ยท benchmarked by Artificial Analysis
Detailed scores
โ€”
Intelligence Index
Overall intelligence across reasoning, knowledge & math evals
โ€”
Coding Index
Coding ability across software-engineering evals
โ€”
Agentic Index
Tool use & multi-step agent task performance
Reasoning
GPQA Diamond 45.2%

Graduate-level questions in biology, chemistry & physics

Humanity's Last Exam 5%

Expert-level questions across many academic domains

IFBench 46.3%

Precise following of detailed instructions

ฯ„ยฒ-Bench Telecom 91.8%

Tool-using agent tasks in a telecom support setting

CritPt 0%

Unpublished physics research reasoning problems

AA-LCR 20.3%

Long-context reasoning across large inputs

Coding
Terminal-Bench Hard 3.8%

Complex command-line and terminal workflows

Knowledge
Omniscience Accuracy 13.1%

Breadth of factual knowledge across domains

Omniscience Non-Hallucination 5.7%

How reliably the model avoids fabricated answers

All variants
Variant GPQA Diamond Humanity's Last Exam IFBench ฯ„ยฒ-Bench Telecom AA-LCR Terminal-Bench Hard CritPt Omniscience Accuracy Omniscience Non-Hallucination
GLM-4.7-Flash (Non-reasoning) 45.2% 5% 46.3% 91.8% 20.3% 3.8% 0% 13.1% 5.7%
GLM-4.7-Flash (Reasoning) 58.1% 7.6% 60.8% 98.8% 41.7% 22% 0.3% 16.2% 6.1%

Design Arena

Elo rating by arena & category
Variant Arena Category Elo Win % Percentile Avg time (ms)
glm-4.7-flash models 3d 1,139 51.2% 42 116,926
glm-4.7-flash models codecategories 1,187 53.1% 55 150,068
glm-4.7-flash models dataviz 1,134 45.3% 30 126,750
glm-4.7-flash models gamedev 1,149 49.7% 41 153,698
glm-4.7-flash models svg 1,046 44.2% 16 81,729
glm-4.7-flash models uicomponent 1,216 57.6% 64 116,336
glm-4.7-flash models website 1,202 54% 58 159,006

Other Models in the GLM 4.7 Family

FAQs about GLM 4.7 FlashX

How much does GLM 4.7 FlashX cost per 1M tokens?

GLM 4.7 FlashX costs $0.07 per million input tokens, and $0.40 per million output tokens.

What does a typical workload cost with GLM 4.7 FlashX?

1,000 requests of 2,000 input and 500 output tokens each โ€” 2,000,000 input and 500,000 output tokens in total โ€” costs $0.34 with GLM 4.7 FlashX at its lowest rates. Use the calculator on this page for your own volumes.

Where can I use GLM 4.7 FlashX?

GLM 4.7 FlashX is available through Z.ai, from $0.07 per million input tokens.

What is GLM 4.7 FlashX's context window?

GLM 4.7 FlashX has a context window of 202,752 tokens, and returns up to 128,000 tokens in a single response.

What input types does GLM 4.7 FlashX support?

GLM 4.7 FlashX accepts text input, and returns text.