Cheaper overall
Gemini 3.8 Flash
11% lower blended rate
Higher benchmark scores
Gemini 3.8 Flash
leads on 11 of 11 figures
Bigger context window
Gemini 3.8 Flash
1,048,576 vs 400,000 tokens
Where each one wins
Gemini 3.8 Flash
- Cheaper output tokens $3.75 vs $4.50, 17% less
- Cheaper on all four workloads
- Larger context window 1,048,576 tokens
- Higher overall intelligence (Intelligence Index) 40.9 vs 24.1
- Higher coding ability (Coding Index) 76.3 vs 56.1
- Higher multi-step tool use (Agentic Index) 40.2 vs 17.9
- Ahead on 8 more benchmark figures
GPT-5.4 Mini
- Longer maximum output 128,000 tokens
- Offers Batch and Flex pricing not listed for Gemini 3.8 Flash
What four workloads cost
One run, and the same run a thousand times.
Gemini 3.8 Flash is cheaper on all four workloads, by much the same margin each time โ about 1.1 times, from $0.0026 against $0.0030 on a chat turn to $0.3113 against $0.3135 on a long document.
| Workload | Gemini 3.8 Flash | GPT-5.4 Mini | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0026 | $0.0030 | $2.63 / $3.00 | Gemini 3.8 Flash is 13% cheaper |
| RAG answer 10,000 in / 800 out | $0.0105 | $0.0111 | $10.50 / $11.10 | Gemini 3.8 Flash is 5% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1275 | $0.1380 | $127.50 / $138.00 | Gemini 3.8 Flash is 8% cheaper |
| Long document 400,000 in / 3,000 out | $0.3113 | $0.3135 | $311.25 / $313.50 | Gemini 3.8 Flash is 0.7% cheaper |
Price per 1M tokens
Neither model changes its rate with prompt length, so these rates apply to every request.
| Context band | Price | Gemini 3.8 Flash | GPT-5.4 Mini | Difference |
|---|---|---|---|---|
| Any prompt length | Input | $0.75 | $0.75 | Same |
| Any prompt length | Cached input | $0.075 | $0.075 | Same |
| Any prompt length | Output | $3.75 | $4.50 | Gemini 3.8 Flash is 17% cheaper |
Blended: Gemini 3.8 Flash $1.50, GPT-5.4 Mini $1.6875 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Gemini 3.8 Flash priced by Perplexity, GPT-5.4 Mini by OpenAI. Standard tier, pay-as-you-go.
Other pricing tiers
Only GPT-5.4 Mini lists Batch and Flex, at up to 50% off its own standard rate; Gemini 3.8 Flash offers none of them.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | Gemini 3.8 Flash | GPT-5.4 Mini | Cheaper |
|---|---|---|---|
| Batch only on GPT-5.4 Mini | Not offered | $0.375 in / $2.25 out 50% off Standard | — |
| Flex only on GPT-5.4 Mini | Not offered | $0.375 in / $2.25 out 50% off Standard | — |
Benchmarks
Gemini 3.8 Flash leads on all 11 figures, furthest ahead on answer reliability (Omniscience Non-Hallucination), by 34.6 points.
Overall intelligence
Intelligence Index โ Overall intelligence across reasoning, knowledge & math evals
Gemini 3.8 Flash ahead by 16.8
Coding ability
Coding Index โ Coding ability across software-engineering evals
Gemini 3.8 Flash ahead by 20.2
Multi-step tool use
Agentic Index โ Tool use & multi-step agent task performance
Gemini 3.8 Flash ahead by 22.3
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
Gemini 3.8 Flash ahead by 7.8
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
Gemini 3.8 Flash ahead by 19.7
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
Gemini 3.8 Flash ahead by 4.3
Scientific coding
SciCode โ Research-level scientific coding tasks
Gemini 3.8 Flash ahead by 4.5
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
Gemini 3.8 Flash ahead by 8.3
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
Gemini 3.8 Flash ahead by 20.6
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
Gemini 3.8 Flash ahead by 17.1
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
Gemini 3.8 Flash ahead by 34.6
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
Gemini 3.8 Flash holds 648,576 more input tokens in one request. Gemini 3.8 Flash also takes audio and video.
| Specification | Gemini 3.8 Flash | GPT-5.4 Mini |
|---|---|---|
| Context window | 1,048,576 tokens | 400,000 tokens |
| Max output | 65,536 tokens | 128,000 tokens |
| Takes in | Text, audio, image, video, file | Text, image, file |
| Puts out | Text | Text |
| Knowledge cutoff | โ | 31 Aug 2025 |
| Released | 2 Sep 2026 | 17 Mar 2026 |
| Status | Active | Active |
| Sold by | Perplexity | OpenAI, Perplexity |
FAQs
- Is Gemini 3.8 Flash cheaper than GPT-5.4 Mini?
- Yes. A 1,000-token prompt with a 500-token reply costs $0.0026 on Gemini 3.8 Flash and $0.0030 on GPT-5.4 Mini, and the same model is cheaper on every workload on this page.
- How much do Gemini 3.8 Flash and GPT-5.4 Mini cost per 1M tokens?
- Gemini 3.8 Flash costs $0.75 for input and $3.75 for output. GPT-5.4 Mini costs $0.75 and $4.50. Standard pay-as-you-go rates, as of 10 Sep 2026.
- Which is better, Gemini 3.8 Flash or GPT-5.4 Mini?
- On published benchmarks, Gemini 3.8 Flash. It leads on 11 of the 11 figures both models report, including overall intelligence (Intelligence Index), where it scores 40.9 against 24.1.
- Does Gemini 3.8 Flash or GPT-5.4 Mini offer batch pricing?
- GPT-5.4 Mini only. Its batch input costs $0.375, 50% off its standard rate, and Gemini 3.8 Flash lists no batch tier at all.
- Does Gemini 3.8 Flash or GPT-5.4 Mini have a bigger context window?
- Gemini 3.8 Flash, at 1,048,576 tokens against 400,000. That is 648,576 more input tokens in a single request.