Cheaper overall
GLM 4.5 Air
6% lower blended rate
Higher benchmark scores
GPT-5.6 Luna
leads on 6 of 6 figures
Bigger context window
GPT-5.6 Luna
1,050,000 vs 131,072 tokens
Cheaper cached input
GPT-5.6 Luna
$0.02 vs $0.03, 33% less
Where each one wins
GPT-5.6 Luna
- Cheaper cached input $0.02 vs $0.03, 33% less
- Larger context window 1,050,000 tokens
- Longer maximum output 128,000 tokens
- Higher graduate-level science (GPQA Diamond) 91.1 vs 73.3
- Higher expert-exam performance (Humanity's Last Exam) 39.5 vs 7
- Higher long-context reasoning (AA-LCR) 83.7 vs 46.7
- Ahead on 3 more benchmark figures
- Offers Batch, Flex and Priority pricing not listed for GLM 4.5 Air
GLM 4.5 Air
- Cheaper output tokens $1.10 vs $1.20, 8% less
What four workloads cost
One run, and the same run a thousand times.
Caching is what decides it: over twenty calls GPT-5.6 Luna costs $0.0368 against $0.0414, a difference of $4.60 across a thousand loops.
| Workload | GPT-5.6 Luna | GLM 4.5 Air | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0008 | $0.0008 | $0.8000 / $0.7500 | GLM 4.5 Air is 6% cheaper |
| RAG answer 10,000 in / 800 out | $0.0030 | $0.0029 | $2.96 / $2.88 | GLM 4.5 Air is 3% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.0368 | $0.0414 | $36.80 / $41.40 | GPT-5.6 Luna is 11% cheaper |
| Long document 400,000 in / 3,000 out | GLM 4.5 Air's context window holds 131,072 tokens, so a 400,000-token prompt does not fit. | |||
Price per 1M tokens
Above 272,000 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: GPT-5.6 Luna's input goes from $0.2000 to $0.4000. It does not change which of the two is cheaper.
| Context band | Price | GPT-5.6 Luna | GLM 4.5 Air | Difference |
|---|---|---|---|---|
| Prompts up to 272,000 tokens | Input | $0.20 | $0.20 | Same |
| Prompts up to 272,000 tokens | Cached input | $0.02 | $0.03 | GPT-5.6 Luna is 33% cheaper |
| Prompts up to 272,000 tokens | Output | $1.20 | $1.10 | GLM 4.5 Air is 8% cheaper |
| Prompts over 272,000 tokens | Input | $0.40 | $0.20 | GLM 4.5 Air is 50% cheaper |
| Prompts over 272,000 tokens | Cached input | $0.04 | $0.03 | GLM 4.5 Air is 25% cheaper |
| Prompts over 272,000 tokens | Output | $1.80 | $1.10 | GLM 4.5 Air is 39% cheaper |
Blended: GPT-5.6 Luna $0.45, GLM 4.5 Air $0.425 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
GPT-5.6 Luna priced by OpenAI, GLM 4.5 Air by Z.ai. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.6 Luna lists Batch, Flex and Priority, at up to 50% off its own standard rate; GLM 4.5 Air offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GPT-5.6 Luna | GLM 4.5 Air | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.6 Luna | $0.10 in / $0.60 out 50% off Standard | Not offered | — |
| Flex only on GPT-5.6 Luna | $0.10 in / $0.60 out 50% off Standard | Not offered | — |
| Priority only on GPT-5.6 Luna | $0.40 in / $2.40 out | Not offered | — |
Benchmarks
GPT-5.6 Luna leads on all 6 figures, furthest ahead on long-context reasoning (AA-LCR), by 37 points.
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GPT-5.6 Luna ahead by 17.8
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GPT-5.6 Luna ahead by 32.5
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.6 Luna ahead by 37
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.6 Luna ahead by 20.6
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GPT-5.6 Luna ahead by 26.4
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GPT-5.6 Luna ahead by 17.9
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.6 Luna holds 918,928 more input tokens in one request. GPT-5.6 Luna also takes file and image. Only GPT-5.6 Luna lists structured outputs and code execution; only GLM 4.5 Air lists open weights.
| Specification | GPT-5.6 Luna | GLM 4.5 Air |
|---|---|---|
| Context window | 1,050,000 tokens | 131,072 tokens |
| Max output | 128,000 tokens | 98,304 tokens |
| Takes in | Text, image, file | Text |
| Puts out | Text | Text |
| Knowledge cutoff | 16 Feb 2026 | 31 Dec 2024 |
| Released | 9 Jul 2026 | 26 Jul 2025 |
| Status | Active | Active |
| Sold by | OpenAI, Perplexity | Z.ai |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Not listed |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Yes | Not listed |
| Computer use | Not listed | Not listed |
| Open weights | Not listed | Yes |
| Reasoning | Yes | Yes |
FAQs
- Is GPT-5.6 Luna cheaper than GLM 4.5 Air?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0008 on GPT-5.6 Luna and $0.0008 on GLM 4.5 Air.
- How much do GPT-5.6 Luna and GLM 4.5 Air cost per 1M tokens?
- GPT-5.6 Luna costs $0.20 for input and $1.20 for output. GLM 4.5 Air costs $0.20 and $1.10. Standard pay-as-you-go rates, as of 30 Jul 2026.
- Which is better, GPT-5.6 Luna or GLM 4.5 Air?
- On published benchmarks, GPT-5.6 Luna. It leads on 6 of the 6 figures both models report, including graduate-level science (GPQA Diamond), where it scores 91.1 against 73.3.
- Does GPT-5.6 Luna or GLM 4.5 Air offer batch pricing?
- GPT-5.6 Luna only. Its batch input costs $0.10, 50% off its standard rate, and GLM 4.5 Air lists no batch tier at all.
- Does GPT-5.6 Luna or GLM 4.5 Air have a bigger context window?
- GPT-5.6 Luna, at 1,050,000 tokens against 131,072. That is 918,928 more input tokens in a single request.