Cheaper overall
GLM-5.3-Flash
47% lower blended rate
Higher benchmark scores
GLM-5.3-Flash
leads on 7 of 11 figures
Bigger context window
gpt-5.6-luna
1,050,000 vs 1,000,000 tokens
Cheaper cached input
gpt-5.6-luna
$0.02 vs $0.03, 33% less
Where each one wins
GLM-5.3-Flash
- Cheaper input tokens $0.15 vs $0.20, 25% less
- Cheaper output tokens $0.50 vs $1.20, 58% less
- Cheaper on all four workloads
- Higher overall intelligence (Intelligence Index) 41.9 vs 37.3
- Higher coding ability (Coding Index) 71.5 vs 71.4
- Higher multi-step tool use (Agentic Index) 51.2 vs 42.1
- Ahead on 4 more benchmark figures
gpt-5.6-luna
- Cheaper cached input $0.02 vs $0.03, 33% less
- Larger context window 1,050,000 tokens
- Higher long-context reasoning (AA-LCR) 83.7 vs 80
- Higher scientific coding (SciCode) 53.6 vs 51.6
- Higher physics research reasoning (CritPt) 20.6 vs 15.4
- Ahead on 1 more benchmark figure
- Offers Batch, Flex and Priority pricing not listed for GLM-5.3-Flash
What four workloads cost
One run, and the same run a thousand times.
The gap opens on long prompts: GLM-5.3-Flash costs $0.0615 against $0.1654, a difference of $103.90 over a thousand runs.
| Workload | GLM-5.3-Flash | gpt-5.6-luna | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0004 | $0.0008 | $0.4000 / $0.8000 | GLM-5.3-Flash is 50% cheaper |
| RAG answer 10,000 in / 800 out | $0.0019 | $0.0030 | $1.90 / $2.96 | GLM-5.3-Flash is 36% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.0310 | $0.0368 | $31.00 / $36.80 | GLM-5.3-Flash is 16% cheaper |
| Long document 400,000 in / 3,000 out | $0.0615 | $0.1654 | $61.50 / $165.40 | GLM-5.3-Flash is 63% cheaper |
Price per 1M tokens
Above 272,000 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: gpt-5.6-luna's input goes from $0.2000 to $0.4000. It does not change which of the two is cheaper.
| Context band | Price | GLM-5.3-Flash | gpt-5.6-luna | Difference |
|---|---|---|---|---|
| Prompts up to 272,000 tokens | Input | $0.15 | $0.20 | GLM-5.3-Flash is 25% cheaper |
| Prompts up to 272,000 tokens | Cached input | $0.03 | $0.02 | gpt-5.6-luna is 33% cheaper |
| Prompts up to 272,000 tokens | Output | $0.50 | $1.20 | GLM-5.3-Flash is 58% cheaper |
| Prompts over 272,000 tokens | Input | $0.15 | $0.40 | GLM-5.3-Flash is 63% cheaper |
| Prompts over 272,000 tokens | Cached input | $0.03 | $0.04 | GLM-5.3-Flash is 25% cheaper |
| Prompts over 272,000 tokens | Output | $0.50 | $1.80 | GLM-5.3-Flash is 72% cheaper |
Blended: GLM-5.3-Flash $0.2375, gpt-5.6-luna $0.45 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
GLM-5.3-Flash priced by Z.ai, gpt-5.6-luna by OpenAI. Standard tier, pay-as-you-go.
Other pricing tiers
Only gpt-5.6-luna lists Batch, Flex and Priority, at up to 50% off its own standard rate; GLM-5.3-Flash offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GLM-5.3-Flash | gpt-5.6-luna | Cheaper |
|---|---|---|---|
| Batch only on gpt-5.6-luna | Not offered | $0.10 in / $0.60 out 50% off Standard | — |
| Flex only on gpt-5.6-luna | Not offered | $0.10 in / $0.60 out 50% off Standard | — |
| Priority only on gpt-5.6-luna | Not offered | $0.40 in / $2.40 out | — |
Benchmarks
They share 11 figures: GLM-5.3-Flash leads on 7, gpt-5.6-luna on 4. The widest gap is 47.4 points, on answer reliability (Omniscience Non-Hallucination).
Overall intelligence
Intelligence Index โ Overall intelligence across reasoning, knowledge & math evals
GLM-5.3-Flash ahead by 4.6
Coding ability
Coding Index โ Coding ability across software-engineering evals
GLM-5.3-Flash ahead by 0.1
Multi-step tool use
Agentic Index โ Tool use & multi-step agent task performance
GLM-5.3-Flash ahead by 9.1
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GLM-5.3-Flash ahead by 0.1
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GLM-5.3-Flash ahead by 0.4
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
gpt-5.6-luna ahead by 3.7
Scientific coding
SciCode โ Research-level scientific coding tasks
gpt-5.6-luna ahead by 2
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
gpt-5.6-luna ahead by 5.2
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
GLM-5.3-Flash ahead by 10.5
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
gpt-5.6-luna ahead by 15.2
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GLM-5.3-Flash ahead by 47.4
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
gpt-5.6-luna holds 50,000 more input tokens in one request. GLM-5.3-Flash also takes video. Only GLM-5.3-Flash lists computer use and open weights.
| Specification | GLM-5.3-Flash | gpt-5.6-luna |
|---|---|---|
| Context window | 1,000,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Takes in | Text, image, video, file | Text, image, file |
| Puts out | Text | Text |
| Knowledge cutoff | โ | 16 Feb 2026 |
| Released | 26 Aug 2026 | 9 Jul 2026 |
| Status | Active | Active |
| Sold by | Perplexity, Z.ai | OpenAI, Perplexity |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Yes |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Yes | Yes |
| Computer use | Yes | Not listed |
| Open weights | Yes | Not listed |
| Reasoning | Yes | Yes |
FAQs
- Is GLM-5.3-Flash cheaper than gpt-5.6-luna?
- Yes. A 1,000-token prompt with a 500-token reply costs $0.0004 on GLM-5.3-Flash and $0.0008 on gpt-5.6-luna, and the same model is cheaper on every workload on this page.
- How much do GLM-5.3-Flash and gpt-5.6-luna cost per 1M tokens?
- GLM-5.3-Flash costs $0.15 for input and $0.50 for output. gpt-5.6-luna costs $0.20 and $1.20. Standard pay-as-you-go rates, as of 26 Aug 2026.
- Which is better, GLM-5.3-Flash or gpt-5.6-luna?
- On published benchmarks, GLM-5.3-Flash. It leads on 7 of the 11 figures both models report, including overall intelligence (Intelligence Index), where it scores 41.9 against 37.3.
- Does GLM-5.3-Flash or gpt-5.6-luna offer batch pricing?
- gpt-5.6-luna only. Its batch input costs $0.10, 50% off its standard rate, and GLM-5.3-Flash lists no batch tier at all.
- Does GLM-5.3-Flash or gpt-5.6-luna have a bigger context window?
- gpt-5.6-luna, at 1,050,000 tokens against 1,000,000. That is 50,000 more input tokens in a single request.