OpenAI reports GPT-5.4 results including 83.0% on GDPval and 87.3% on SpreadsheetBench, with long-context and tool-search boundaries.
OpenAI Newsroom · Read evidenceGPT-5.4 · Reviews and evidence
Which GPT-5.4 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
With one one-shot atomic-clock prompt, GPT-5.4 looked best but synchronization drifted; the article calls this a single-task observation.
Thomas Wiegold Blog · Read evidenceA four-day Reddit discussion reports GPT-5.4 helping with review and multi-step execution, alongside forgotten constraints, excessive calls, and xhigh cost.
Reddit r/AIAgents · Read evidenceFull reviews and related reading
Selected evidence
GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark
OpenAI reports GPT-5.4 results including 83.0% on GDPval and 87.3% on SpreadsheetBench, with long-context and tool-search boundaries.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
GPT-5.4: A Four-Model Comparison of Atomic Clock Applications
With one one-shot atomic-clock prompt, GPT-5.4 looked best but synchronization drifted; the article calls this a single-task observation.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
GPT-5.4: Reddit AI Agents — Multi-step Agents and Model Routing Experience
A four-day Reddit discussion reports GPT-5.4 helping with review and multi-step execution, alongside forgotten constraints, excessive calls, and xhigh cost.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
All sources
All sources
GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark
OpenAI reports GPT-5.4 results including 83.0% on GDPval and 87.3% on SpreadsheetBench, with long-context and tool-search boundaries.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
GPT-5.4: A Four-Model Comparison of Atomic Clock Applications
With one one-shot atomic-clock prompt, GPT-5.4 looked best but synchronization drifted; the article calls this a single-task observation.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
GPT-5.4: Reddit AI Agents — Multi-step Agents and Model Routing Experience
A four-day Reddit discussion reports GPT-5.4 helping with review and multi-step execution, alongside forgotten constraints, excessive calls, and xhigh cost.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
GPT-5.4
Compare GPT-5.4 in Tabbit
Model access, features, and permissions depend on your current client account.