GLM 4.6 Pricing - Cost Calculator

GLM-4.6 is Z.aiโ€™s open-weight reasoning and coding model built for long-context software engineering, tool use, autonomous agents, and complex technical workflows.

Released 30 September 2025, GLM 4.6 is a coding model from Z.ai. It pairs a 204,800-token context window with up to 16,384 tokens of output, accepting text and returning text. Training data ends 31 March 2025, and the model reasons step by step. It costs $0.60 per million input tokens through Z.ai.

Last updated Sep 30, 2026

Never miss an AI pricing or launch update

Get notified about new model launches, pricing changes, provider updates, deprecations, and cost-saving insights.

๐Ÿ”’ We respect your privacy. Unsubscribe anytime.

Model Details

Released
Sep 30, 2025
Knowledge Cutoff
Mar 31, 2025
Context Length
204,800
Max Output
16,384
Modalities
Text → Text
Capabilities
Tool use Structured outputs Prompt caching Open weights

Where is GLM 4.6 a perfect fit?

GLM-4.6 expands context to 200K tokens while improving coding, reasoning, tool use, search agents, frontend generation, writing quality, and real-world multi-step agentic performance. It works well for below use-cases:
- Software engineering โ€” repository-level coding, debugging, testing, and refactoring
- Coding agents โ€” Claude Code, Cline, Roo Code, Kilo Code and similar frameworks
- Tool-using agents โ€” multi-step workflows involving external tools and APIs
- Search agents โ€” research and information-retrieval workflows
- Frontend development โ€” polished webpages and UI generation
- Long-context analysis โ€” large codebases and lengthy technical documents
- Complex reasoning โ€” technical analysis and multi-step problem solving

GLM-4.6 was a major shift toward agentic software engineering for the GLM family, increasing the context window from 128K to 200K and emphasizing coding, tool use, search agents, and real-world multi-step tasks.

Quick Model Estimate

(USD 0.6000 per 1M tokens)
(USD 2.2000 per 1M tokens)

Your GLM 4.6 Cost Estimate

๐Ÿ’ฐ Total Cost

โ€”

for 1000 input + 1000 output tokens

๐Ÿ“ฅ Input (1000 ร— $0.600000) โ€”
๐Ÿ“ค Output (1000 ร— $2.200000) โ€”

Cost Breakdown

๐Ÿ“ฅ Input ๐Ÿ“ค Output

Prices are watched for changes and checked against the provider's own page.

Pricing

Provider โ†•
Modality โ†•
Service Tier โ†•
Input Price
(per 1M tokens)
โ†•
Output Price
(per 1M tokens)
โ†•
Cached Input
(per 1M tokens)
โ†•
Context Size โ†•
View
Z.ai LogoZ.aiTextStandard$0.6000$2.2000$0.1100204,800 tokensโ†’

Benchmarks

Scores from standardized evaluations by Artificial Analysis and Design Arena. Higher is better โ€” the indices summarize overall ability, while the detailed scores break down performance on individual benchmarks.

Artificial Analysis

Higher is better ยท benchmarked by Artificial Analysis
Detailed scores
โ€”
Intelligence Index
Overall intelligence across reasoning, knowledge & math evals
โ€”
Coding Index
Coding ability across software-engineering evals
โ€”
Agentic Index
Tool use & multi-step agent task performance
Reasoning
GPQA Diamond 63.2%

Graduate-level questions in biology, chemistry & physics

Humanity's Last Exam 5.5%

Expert-level questions across many academic domains

IFBench 36.7%

Precise following of detailed instructions

ฯ„ยฒ-Bench Telecom 76.9%

Tool-using agent tasks in a telecom support setting

CritPt 0%

Unpublished physics research reasoning problems

AA-LCR 26.3%

Long-context reasoning across large inputs

Coding
Terminal-Bench Hard 28.8%

Complex command-line and terminal workflows

Knowledge
Omniscience Accuracy 21.4%

Breadth of factual knowledge across domains

Omniscience Non-Hallucination 32.4%

How reliably the model avoids fabricated answers

All variants
Variant GPQA Diamond Humanity's Last Exam IFBench ฯ„ยฒ-Bench Telecom AA-LCR Terminal-Bench Hard CritPt GDPval Omniscience Accuracy Omniscience Non-Hallucination
GLM-4.6 (Non-reasoning) 63.2% 5.5% 36.7% 76.9% 26.3% 28.8% 0% โ€” 21.4% 32.4%
GLM-4.6 (Reasoning) 78% 14.5% 43.4% 70.5% 54% 25% 1.1% 12.3% 26.9% 5.9%

Design Arena

Elo rating by arena & category
Variant Arena Category Elo Win % Percentile Avg time (ms)
glm-4.6 agents androidnative 1,099 52.6% 19 โ€”
glm-4.6 agents fullstack 1,028 42.3% 11 โ€”
glm-4.6 agents godotgamedev 1,180 53.2% 53 โ€”
glm-4.6 agents mobileapps 1,124 47.2% 20 โ€”
glm-4.6 models 3d 1,147 54.2% 46 80,173
glm-4.6 models codecategories 1,175 54.3% 48 109,303
glm-4.6 models dataviz 1,171 52.6% 43 86,069
glm-4.6 models gamedev 1,165 54.8% 47 108,718
glm-4.6 models svg 1,117 52.1% 37 33,248
glm-4.6 models uicomponent 1,156 52.5% 42 86,789
glm-4.6 models website 1,182 54.3% 47 115,091

FAQs about GLM 4.6

How much does GLM 4.6 cost per 1M tokens?

GLM 4.6 costs $0.60 per million input tokens, and $2.20 per million output tokens.

What does a typical workload cost with GLM 4.6?

1,000 requests of 2,000 input and 500 output tokens each โ€” 2,000,000 input and 500,000 output tokens in total โ€” costs $2.30 with GLM 4.6 at its lowest rates. Use the calculator on this page for your own volumes.

Where can I use GLM 4.6?

GLM 4.6 is available through Z.ai, from $0.60 per million input tokens.

What is GLM 4.6's context window?

GLM 4.6 has a context window of 204,800 tokens, and returns up to 16,384 tokens in a single response.

What input types does GLM 4.6 support?

GLM 4.6 accepts text input, and returns text.

What is GLM 4.6's knowledge cutoff?

GLM 4.6's training data runs to 31 March 2025. It has no built-in knowledge of events after that date.