Cheaper overall
Gemini 3.8 Flash
73% lower blended rate
Higher benchmark scores
Gemini 3.8 Flash
leads on 17 of 19 figures
Bigger context window
GPT-5.4
1,050,000 vs 1,048,576 tokens
Cheaper cached input
Gemini 3.8 Flash
$0.075 vs $0.25, 70% less
Where each one wins
Gemini 3.8 Flash
- Cheaper input tokens $0.75 vs $2.50, 70% less
- Cheaper output tokens $3.75 vs $15.00, 75% less
- Cheaper cached input $0.075 vs $0.25, 70% less
- Cheaper on all four workloads
- Higher coding ability (Coding Index) 76.3 vs 71.1
- Higher graduate-level science (GPQA Diamond) 95.3 vs 92
- Higher expert-exam performance (Humanity's Last Exam) 47.8 vs 43.7
- Ahead on 14 more benchmark figures
GPT-5.4
- Larger context window 1,050,000 tokens
- Longer maximum output 128,000 tokens
- Higher long-context reasoning (AA-LCR) 82 vs 81.3
- Higher physics research reasoning (CritPt) 23.4 vs 18.3
- Offers Batch, Flex and Priority pricing not listed for Gemini 3.8 Flash
What four workloads cost
One run, and the same run a thousand times.
Gemini 3.8 Flash is cheaper on all four workloads, by much the same margin each time โ about 4.4 times, from $0.0026 against $0.0100 on a chat turn to $0.3113 against $2.07 on a long document.
| Workload | Gemini 3.8 Flash | GPT-5.4 | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0026 | $0.0100 | $2.63 / $10.00 | Gemini 3.8 Flash is 74% cheaper |
| RAG answer 10,000 in / 800 out | $0.0105 | $0.0370 | $10.50 / $37.00 | Gemini 3.8 Flash is 72% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1275 | $0.4600 | $127.50 / $460.00 | Gemini 3.8 Flash is 72% cheaper |
| Long document 400,000 in / 3,000 out | $0.3113 | $2.07 | $311.25 / $2,067.50 | Gemini 3.8 Flash is 85% cheaper |
Price per 1M tokens
Above 272,000 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: GPT-5.4's input goes from $2.50 to $5.00. It does not change which of the two is cheaper.
| Context band | Price | Gemini 3.8 Flash | GPT-5.4 | Difference |
|---|---|---|---|---|
| Prompts up to 272,000 tokens | Input | $0.75 | $2.50 | Gemini 3.8 Flash is 70% cheaper |
| Prompts up to 272,000 tokens | Cached input | $0.075 | $0.25 | Gemini 3.8 Flash is 70% cheaper |
| Prompts up to 272,000 tokens | Output | $3.75 | $15.00 | Gemini 3.8 Flash is 75% cheaper |
| Prompts over 272,000 tokens | Input | $0.75 | $5.00 | Gemini 3.8 Flash is 85% cheaper |
| Prompts over 272,000 tokens | Cached input | $0.075 | $0.50 | Gemini 3.8 Flash is 85% cheaper |
| Prompts over 272,000 tokens | Output | $3.75 | $22.50 | Gemini 3.8 Flash is 83% cheaper |
Blended: Gemini 3.8 Flash $1.50, GPT-5.4 $5.625 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Gemini 3.8 Flash priced by Perplexity, GPT-5.4 by OpenAI. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.4 lists Batch, Flex and Priority, at up to 50% off its own standard rate; Gemini 3.8 Flash offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | Gemini 3.8 Flash | GPT-5.4 | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.4 | Not offered | $1.25 in / $7.50 out 50% off Standard | — |
| Flex only on GPT-5.4 | Not offered | $1.25 in / $7.50 out 50% off Standard | — |
| Priority only on GPT-5.4 | Not offered | $5.00 in / $30.00 out | — |
Benchmarks
They share 19 figures: Gemini 3.8 Flash leads on 17, GPT-5.4 on 2. The widest gap is 227 points, on Fullstack Elo.
Coding ability
Coding Index โ Coding ability across software-engineering evals
Gemini 3.8 Flash ahead by 5.2
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
Gemini 3.8 Flash ahead by 3.3
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
Gemini 3.8 Flash ahead by 4.1
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.4 ahead by 0.7
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.4 ahead by 5.1
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
Gemini 3.8 Flash ahead by 9
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
Gemini 3.8 Flash ahead by 3.8
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
Gemini 3.8 Flash ahead by 27.4
Androidnative Elo
Design Arena head-to-head rating in the androidnative category
Gemini 3.8 Flash ahead by 58
Fullstack Elo
Design Arena head-to-head rating in the fullstack category
Gemini 3.8 Flash ahead by 227
Godotgamedev Elo
Design Arena head-to-head rating in the godotgamedev category
Gemini 3.8 Flash ahead by 138
Mobileapps Elo
Design Arena head-to-head rating in the mobileapps category
Gemini 3.8 Flash ahead by 143
Webapps Elo
Design Arena head-to-head rating in the webapps category
Gemini 3.8 Flash ahead by 182
3d Elo
Design Arena head-to-head rating in the 3d category
Gemini 3.8 Flash ahead by 196
Codecategories Elo
Design Arena head-to-head rating in the codecategories category
Gemini 3.8 Flash ahead by 94
Dataviz Elo
Design Arena head-to-head rating in the dataviz category
Gemini 3.8 Flash ahead by 27
Gamedev Elo
Design Arena head-to-head rating in the gamedev category
Gemini 3.8 Flash ahead by 73
Uicomponent Elo
Design Arena head-to-head rating in the uicomponent category
Gemini 3.8 Flash ahead by 80
Website Elo
Design Arena head-to-head rating in the website category
Gemini 3.8 Flash ahead by 79
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.4 holds 1,424 more input tokens in one request. Gemini 3.8 Flash also takes audio and video.
| Specification | Gemini 3.8 Flash | GPT-5.4 |
|---|---|---|
| Context window | 1,048,576 tokens | 1,050,000 tokens |
| Max output | 65,536 tokens | 128,000 tokens |
| Takes in | Text, audio, image, video, file | Text, image, file |
| Puts out | Text | Text |
| Knowledge cutoff | โ | 31 Aug 2025 |
| Released | 2 Sep 2026 | 5 Mar 2026 |
| Status | Active | Active |
| Sold by | Perplexity | OpenAI, Perplexity |
FAQs
- Is Gemini 3.8 Flash cheaper than GPT-5.4?
- Yes. A 1,000-token prompt with a 500-token reply costs $0.0026 on Gemini 3.8 Flash and $0.0100 on GPT-5.4, and the same model is cheaper on every workload on this page.
- How much do Gemini 3.8 Flash and GPT-5.4 cost per 1M tokens?
- Gemini 3.8 Flash costs $0.75 for input and $3.75 for output. GPT-5.4 costs $2.50 and $15.00. Standard pay-as-you-go rates, as of 10 Sep 2026.
- Which is better, Gemini 3.8 Flash or GPT-5.4?
- On published benchmarks, Gemini 3.8 Flash. It leads on 17 of the 19 figures both models report, including coding ability (Coding Index), where it scores 76.3 against 71.1.
- Does Gemini 3.8 Flash or GPT-5.4 offer batch pricing?
- GPT-5.4 only. Its batch input costs $1.25, 50% off its standard rate, and Gemini 3.8 Flash lists no batch tier at all.
- Does Gemini 3.8 Flash or GPT-5.4 have a bigger context window?
- GPT-5.4, at 1,050,000 tokens against 1,048,576. That is 1,424 more input tokens in a single request.