Cheaper overall
GLM 4.7 FlashX
66% lower blended rate
Higher benchmark scores
GPT-5.6 Luna
leads on 6 of 6 figures
Bigger context window
GPT-5.6 Luna
1,050,000 vs 202,752 tokens
Cheaper cached input
GLM 4.7 FlashX
$0.01 vs $0.02, 50% less
Where each one wins
GPT-5.6 Luna
- Larger context window 1,050,000 tokens
- Higher graduate-level science (GPQA Diamond) 91.1 vs 58.1
- Higher expert-exam performance (Humanity's Last Exam) 39.5 vs 7.6
- Higher long-context reasoning (AA-LCR) 83.7 vs 41.7
- Ahead on 3 more benchmark figures
- Offers Batch, Flex and Priority pricing not listed for GLM 4.7 FlashX
GLM 4.7 FlashX
- Cheaper input tokens $0.07 vs $0.20, 65% less
- Cheaper output tokens $0.40 vs $1.20, 67% less
- Cheaper cached input $0.01 vs $0.02, 50% less
- Cheaper on all four workloads
What four workloads cost
One run, and the same run a thousand times.
GLM 4.7 FlashX is cheaper on all four workloads, by much the same margin each time โ about 2.8 times, from $0.0003 against $0.0008 on a chat turn to $0.0144 against $0.0368 on a long document.
| Workload | GPT-5.6 Luna | GLM 4.7 FlashX | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0008 | $0.0003 | $0.8000 / $0.2700 | GLM 4.7 FlashX is 66% cheaper |
| RAG answer 10,000 in / 800 out | $0.0030 | $0.0010 | $2.96 / $1.02 | GLM 4.7 FlashX is 66% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.0368 | $0.0144 | $36.80 / $14.40 | GLM 4.7 FlashX is 61% cheaper |
| Long document 400,000 in / 3,000 out | GLM 4.7 FlashX's context window holds 202,752 tokens, so a 400,000-token prompt does not fit. | |||
Price per 1M tokens
Above 272,000 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: GPT-5.6 Luna's input goes from $0.2000 to $0.4000. It does not change which of the two is cheaper.
| Context band | Price | GPT-5.6 Luna | GLM 4.7 FlashX | Difference |
|---|---|---|---|---|
| Prompts up to 272,000 tokens | Input | $0.20 | $0.07 | GLM 4.7 FlashX is 65% cheaper |
| Prompts up to 272,000 tokens | Cached input | $0.02 | $0.01 | GLM 4.7 FlashX is 50% cheaper |
| Prompts up to 272,000 tokens | Output | $1.20 | $0.40 | GLM 4.7 FlashX is 67% cheaper |
| Prompts over 272,000 tokens | Input | $0.40 | $0.07 | GLM 4.7 FlashX is 83% cheaper |
| Prompts over 272,000 tokens | Cached input | $0.04 | $0.01 | GLM 4.7 FlashX is 75% cheaper |
| Prompts over 272,000 tokens | Output | $1.80 | $0.40 | GLM 4.7 FlashX is 78% cheaper |
Blended: GPT-5.6 Luna $0.45, GLM 4.7 FlashX $0.1525 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
GPT-5.6 Luna priced by OpenAI, GLM 4.7 FlashX by Z.ai. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.6 Luna lists Batch, Flex and Priority, at up to 50% off its own standard rate; GLM 4.7 FlashX offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GPT-5.6 Luna | GLM 4.7 FlashX | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.6 Luna | $0.10 in / $0.60 out 50% off Standard | Not offered | — |
| Flex only on GPT-5.6 Luna | $0.10 in / $0.60 out 50% off Standard | Not offered | — |
| Priority only on GPT-5.6 Luna | $0.40 in / $2.40 out | Not offered | — |
Benchmarks
GPT-5.6 Luna leads on all 6 figures, furthest ahead on long-context reasoning (AA-LCR), by 42 points.
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GPT-5.6 Luna ahead by 33
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GPT-5.6 Luna ahead by 31.9
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.6 Luna ahead by 42
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.6 Luna ahead by 20.3
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GPT-5.6 Luna ahead by 26.5
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GPT-5.6 Luna ahead by 18.9
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.6 Luna holds 847,248 more input tokens in one request. GPT-5.6 Luna also takes file and image. Only GPT-5.6 Luna lists code execution; only GLM 4.7 FlashX lists open weights.
| Specification | GPT-5.6 Luna | GLM 4.7 FlashX |
|---|---|---|
| Context window | 1,050,000 tokens | 202,752 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Takes in | Text, image, file | Text |
| Puts out | Text | Text |
| Knowledge cutoff | 16 Feb 2026 | โ |
| Released | 9 Jul 2026 | 19 Jan 2026 |
| Status | Active | Active |
| Sold by | OpenAI, Perplexity | Z.ai |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Yes |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Yes | Not listed |
| Computer use | Not listed | Not listed |
| Open weights | Not listed | Yes |
| Reasoning | Yes | Yes |
FAQs
- Is GPT-5.6 Luna cheaper than GLM 4.7 FlashX?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0008 on GPT-5.6 Luna and $0.0003 on GLM 4.7 FlashX, and the same model is cheaper on every workload on this page.
- How much do GPT-5.6 Luna and GLM 4.7 FlashX cost per 1M tokens?
- GPT-5.6 Luna costs $0.20 for input and $1.20 for output. GLM 4.7 FlashX costs $0.07 and $0.40. Standard pay-as-you-go rates, as of 30 Jul 2026.
- Which is better, GPT-5.6 Luna or GLM 4.7 FlashX?
- On published benchmarks, GPT-5.6 Luna. It leads on 6 of the 6 figures both models report, including graduate-level science (GPQA Diamond), where it scores 91.1 against 58.1.
- Does GPT-5.6 Luna or GLM 4.7 FlashX offer batch pricing?
- GPT-5.6 Luna only. Its batch input costs $0.10, 50% off its standard rate, and GLM 4.7 FlashX lists no batch tier at all.
- Does GPT-5.6 Luna or GLM 4.7 FlashX have a bigger context window?
- GPT-5.6 Luna, at 1,050,000 tokens against 202,752. That is 847,248 more input tokens in a single request.