Cheaper overall
GPT-5.4 Mini
70% lower blended rate
Higher benchmark scores
GPT-5.4
leads on 11 of 11 figures
Bigger context window
GPT-5.4
1,050,000 vs 400,000 tokens
Cheaper cached input
GPT-5.4 Mini
$0.075 vs $0.25, 70% less
Where each one wins
GPT-5.4 Mini
- Cheaper input tokens $0.75 vs $2.50, 70% less
- Cheaper output tokens $4.50 vs $15.00, 70% less
- Cheaper cached input $0.075 vs $0.25, 70% less
- Cheaper on all four workloads
GPT-5.4
- Larger context window 1,050,000 tokens
- Higher coding ability (Coding Index) 71.1 vs 56.1
- Higher graduate-level science (GPQA Diamond) 92 vs 87.5
- Higher expert-exam performance (Humanity's Last Exam) 43.7 vs 28.1
- Ahead on 8 more benchmark figures
- Offers Priority pricing not listed for GPT-5.4 Mini
What four workloads cost
One run, and the same run a thousand times.
GPT-5.4 Mini is cheaper on all four workloads, by much the same margin each time โ about 4.1 times, from $0.0030 against $0.0100 on a chat turn to $0.3135 against $2.07 on a long document.
| Workload | GPT-5.4 Mini | GPT-5.4 | ร1,000 | Difference |
|---|---|---|---|---|
| Chat turn 1,000 in / 500 out | $0.0030 | $0.0100 | $3.00 / $10.00 | GPT-5.4 Mini is 70% cheaper |
| RAG answer 10,000 in / 800 out | $0.0111 | $0.0370 | $11.10 / $37.00 | GPT-5.4 Mini is 70% cheaper |
| Agent loop 32,000 in / 700 out ร 20 calls | $0.1380 | $0.4600 | $138.00 / $460.00 | GPT-5.4 Mini is 70% cheaper |
| Long document 400,000 in / 3,000 out | $0.3135 | $2.07 | $313.50 / $2,067.50 | GPT-5.4 Mini is 85% cheaper |
Price per 1M tokens
Above 272,000 input tokens the rates change, and the new rate applies to the whole request, not just the tokens past that point: GPT-5.4's input goes from $2.50 to $5.00. It does not change which of the two is cheaper.
| Context band | Price | GPT-5.4 Mini | GPT-5.4 | Difference |
|---|---|---|---|---|
| Prompts up to 272,000 tokens | Input | $0.75 | $2.50 | GPT-5.4 Mini is 70% cheaper |
| Prompts up to 272,000 tokens | Cached input | $0.075 | $0.25 | GPT-5.4 Mini is 70% cheaper |
| Prompts up to 272,000 tokens | Output | $4.50 | $15.00 | GPT-5.4 Mini is 70% cheaper |
| Prompts over 272,000 tokens | Input | $0.75 | $5.00 | GPT-5.4 Mini is 85% cheaper |
| Prompts over 272,000 tokens | Cached input | $0.075 | $0.50 | GPT-5.4 Mini is 85% cheaper |
| Prompts over 272,000 tokens | Output | $4.50 | $22.50 | GPT-5.4 Mini is 80% cheaper |
Blended: GPT-5.4 Mini $1.6875, GPT-5.4 $5.625 per 1M โ one rate at 3:1 input to output, for comparing two models at a glance.
Both priced by OpenAI. Standard tier, pay-as-you-go.
Other pricing tiers
Both models offer Batch and Flex, worth up to 50% off their standard rates, and the cheaper model is the same one on every tier โ so the tier you buy changes the bill rather than the choice. Only GPT-5.4 lists Priority; GPT-5.4 Mini does not offer it.
Every tier either model sells is here, so this is the whole price sheet. A saving is measured against that model's own standard rate โ how we read tiers.
| Tier | GPT-5.4 Mini | GPT-5.4 | Cheaper |
|---|---|---|---|
| Batch | $0.375 in / $2.25 out 50% off Standard | $1.25 in / $7.50 out 50% off Standard | GPT-5.4 Mini |
| Flex | $0.375 in / $2.25 out 50% off Standard | $1.25 in / $7.50 out 50% off Standard | GPT-5.4 Mini |
| Priority only on GPT-5.4 | Not offered | $5.00 in / $30.00 out | — |
Benchmarks
GPT-5.4 leads on all 11 figures, furthest ahead on expert-exam performance (Humanity's Last Exam), by 15.6 points.
Coding ability
Coding Index โ Coding ability across software-engineering evals
GPT-5.4 ahead by 15
Graduate-level science
GPQA Diamond โ Graduate-level questions in biology, chemistry & physics
GPT-5.4 ahead by 4.5
Expert-exam performance
Humanity's Last Exam โ Expert-level questions across many academic domains
GPT-5.4 ahead by 15.6
Instruction following
IFBench โ Precise following of detailed instructions
GPT-5.4 ahead by 0.6
Support-agent performance
ฯยฒ-Bench Telecom โ Tool-using agent tasks in a telecom support setting
GPT-5.4 ahead by 3.8
Long-context reasoning
AA-LCR โ Long-context reasoning across large inputs
GPT-5.4 ahead by 5
Command-line work
Terminal-Bench Hard โ Complex command-line and terminal workflows
GPT-5.4 ahead by 5.3
Physics research reasoning
CritPt โ Unpublished physics research reasoning problems
GPT-5.4 ahead by 13.4
Real-world task performance
GDPval โ Economically valuable, real-world knowledge work
GPT-5.4 ahead by 11.6
Factual accuracy
Omniscience Accuracy โ Breadth of factual knowledge across domains
GPT-5.4 ahead by 13.3
Answer reliability
Omniscience Non-Hallucination โ How reliably the model avoids fabricated answers
GPT-5.4 ahead by 7.2
Each figure is that model's best published run, with the effort level named beside it โ where these scores come from.
Specifications
GPT-5.4 holds 650,000 more input tokens in one request.
| Specification | GPT-5.4 Mini | GPT-5.4 |
|---|---|---|
| Context window | 400,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Takes in | Text, image, file | Text, image, file |
| Puts out | Text | Text |
| Knowledge cutoff | 31 Aug 2025 | 31 Aug 2025 |
| Released | 17 Mar 2026 | 5 Mar 2026 |
| Status | Active | Active |
| Sold by | OpenAI, Perplexity | OpenAI, Perplexity |
FAQs
- Is GPT-5.4 Mini cheaper than GPT-5.4?
- Yes. A 1,000-token prompt with a 500-token reply costs $0.0030 on GPT-5.4 Mini and $0.0100 on GPT-5.4, and the same model is cheaper on every workload on this page.
- How much do GPT-5.4 Mini and GPT-5.4 cost per 1M tokens?
- GPT-5.4 Mini costs $0.75 for input and $4.50 for output. GPT-5.4 costs $2.50 and $15.00. Standard pay-as-you-go rates.
- Which is better, GPT-5.4 Mini or GPT-5.4?
- On published benchmarks, GPT-5.4. It leads on 11 of the 11 figures both models report, including coding ability (Coding Index), where it scores 71.1 against 56.1.
- Is Batch pricing cheaper for GPT-5.4 Mini or GPT-5.4?
- Batch input costs $0.375 on GPT-5.4 Mini and $1.25 on GPT-5.4, 50% below the standard rate on both.
- Does GPT-5.4 Mini or GPT-5.4 have a bigger context window?
- GPT-5.4, at 1,050,000 tokens against 400,000. That is 650,000 more input tokens in a single request.