Cheaper overall
GPT-5.6 Luna
70% lower blended rate
Higher benchmark scores
Gemini 3.8 Flash
leads on 7 of 11 figures
Bigger context window
GPT-5.6 Luna
1,050,000 vs 1,048,576 tokens
Cheaper cached input
GPT-5.6 Luna
$0.02 vs $0.075, 73% less
Where each one wins
Gemini 3.8 Flash
- Higher overall intelligence (Intelligence Index) 40.9 vs 37.3
- Higher coding ability (Coding Index) 76.3 vs 71.4
- Higher graduate-level science (GPQA Diamond) 95.3 vs 91.1
- Ahead on 4 more benchmark figures
GPT-5.6 Luna
- Cheaper input tokens $0.20 vs $0.75, 73% less
- Cheaper output tokens $1.20 vs $3.75, 68% less
- Cheaper cached input $0.02 vs $0.075, 73% less
- Cheaper on all four workloads
- Larger context window 1,050,000 tokens
- Longer maximum output 128,000 tokens
- Higher multi-step tool use (Agentic Index) 42.1 vs 40.2
- Higher long-context reasoning (AA-LCR) 83.7 vs 81.3
- Higher physics research reasoning (CritPt) 20.6 vs 18.3
- Ahead on 1 more benchmark figure
- Offers Batch, Flex and Priority pricing not listed for Gemini 3.8 Flash
What four workloads cost
One run, and the same run a thousand times.
GPT-5.6 Luna is cheaper on all four workloads, by much the same margin each time โ about 3 times, from $0.0008 against $0.0026 on a chat turn to $0.1654 against $0.3113 on a long document.
| Workload | Gemini 3.8 Flash | GPT-5.6 Luna | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0026 | $0.0008 | $2.63 / $0.8000 | GPT-5.6 Luna is 70% cheaper |
| RAG answer 10,000 in / 800 out | $0.0105 | $0.0030 | $10.50 / $2.96 | GPT-5.6 Luna is 72% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1275 | $0.0368 | $127.50 / $36.80 | GPT-5.6 Luna is 71% cheaper |
| Long document 400,000 in / 3,000 out | $0.3113 | $0.1654 | $311.25 / $165.40 | GPT-5.6 Luna is 47% cheaper |
Price per 1M tokens
Above 272,000 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: GPT-5.6 Luna's input goes from $0.2000 to $0.4000. It does not change which of the two is cheaper.
| Context band | Price | Gemini 3.8 Flash | GPT-5.6 Luna | Difference |
|---|---|---|---|---|
| Prompts up to 272,000 tokens | Input | $0.75 | $0.20 | GPT-5.6 Luna is 73% cheaper |
| Prompts up to 272,000 tokens | Cached input | $0.075 | $0.02 | GPT-5.6 Luna is 73% cheaper |
| Prompts up to 272,000 tokens | Output | $3.75 | $1.20 | GPT-5.6 Luna is 68% cheaper |
| Prompts over 272,000 tokens | Input | $0.75 | $0.40 | GPT-5.6 Luna is 47% cheaper |
| Prompts over 272,000 tokens | Cached input | $0.075 | $0.04 | GPT-5.6 Luna is 47% cheaper |
| Prompts over 272,000 tokens | Output | $3.75 | $1.80 | GPT-5.6 Luna is 52% cheaper |
Blended: Gemini 3.8 Flash $1.50, GPT-5.6 Luna $0.45 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Gemini 3.8 Flash priced by Perplexity, GPT-5.6 Luna by OpenAI. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.6 Luna lists Batch, Flex and Priority, at up to 50% off its own standard rate; Gemini 3.8 Flash offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | Gemini 3.8 Flash | GPT-5.6 Luna | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.6 Luna | Not offered | $0.10 in / $0.60 out 50% off Standard | — |
| Flex only on GPT-5.6 Luna | Not offered | $0.10 in / $0.60 out 50% off Standard | — |
| Priority only on GPT-5.6 Luna | Not offered | $0.40 in / $2.40 out | — |
Benchmarks
They share 11 figures: Gemini 3.8 Flash leads on 7, GPT-5.6 Luna on 4. The widest gap is 19.8 points, on answer reliability (Omniscience Non-Hallucination).
Overall intelligence
Intelligence Index โ Overall intelligence across reasoning, knowledge & math evals
Gemini 3.8 Flash ahead by 3.6
Coding ability
Coding Index โ Coding ability across software-engineering evals
Gemini 3.8 Flash ahead by 4.9
Multi-step tool use
Agentic Index โ Tool use & multi-step agent task performance
GPT-5.6 Luna ahead by 1.9
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
Gemini 3.8 Flash ahead by 4.2
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
Gemini 3.8 Flash ahead by 8.3
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.6 Luna ahead by 2.4
Scientific coding
SciCode โ Research-level scientific coding tasks
Gemini 3.8 Flash ahead by 3
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.6 Luna ahead by 2.3
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
GPT-5.6 Luna ahead by 1.6
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
Gemini 3.8 Flash ahead by 11.9
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
Gemini 3.8 Flash ahead by 19.8
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.6 Luna holds 1,424 more input tokens in one request. Gemini 3.8 Flash also takes audio and video. Only Gemini 3.8 Flash lists computer use.
| Specification | Gemini 3.8 Flash | GPT-5.6 Luna |
|---|---|---|
| Context window | 1,048,576 tokens | 1,050,000 tokens |
| Max output | 65,536 tokens | 128,000 tokens |
| Takes in | Text, audio, image, video, file | Text, image, file |
| Puts out | Text | Text |
| Knowledge cutoff | โ | 16 Feb 2026 |
| Released | 2 Sep 2026 | 9 Jul 2026 |
| Status | Active | Active |
| Sold by | Perplexity | OpenAI, Perplexity |
| Tool use | Yes | Yes |
| Structured outputs | Yes | Yes |
| Web search | Not listed | Not listed |
| Prompt caching | Yes | Yes |
| Code execution | Yes | Yes |
| Computer use | Yes | Not listed |
| Open weights | Not listed | Not listed |
| Reasoning | Yes | Yes |
FAQs
- Is Gemini 3.8 Flash cheaper than GPT-5.6 Luna?
- No โ the other way round. A 1,000-token prompt with a 500-token reply costs $0.0026 on Gemini 3.8 Flash and $0.0008 on GPT-5.6 Luna, and the same model is cheaper on every workload on this page.
- How much do Gemini 3.8 Flash and GPT-5.6 Luna cost per 1M tokens?
- Gemini 3.8 Flash costs $0.75 for input and $3.75 for output. GPT-5.6 Luna costs $0.20 and $1.20. Standard pay-as-you-go rates, as of 10 Sep 2026.
- Which is better, Gemini 3.8 Flash or GPT-5.6 Luna?
- On published benchmarks, Gemini 3.8 Flash. It leads on 7 of the 11 figures both models report, including overall intelligence (Intelligence Index), where it scores 40.9 against 37.3.
- Does Gemini 3.8 Flash or GPT-5.6 Luna offer batch pricing?
- GPT-5.6 Luna only. Its batch input costs $0.10, 50% off its standard rate, and Gemini 3.8 Flash lists no batch tier at all.
- Does Gemini 3.8 Flash or GPT-5.6 Luna have a bigger context window?
- GPT-5.6 Luna, at 1,050,000 tokens against 1,048,576. That is 1,424 more input tokens in a single request.