OpenAI NewsroomVendor report
GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark
OpenAI reports GPT-5.4 results including 83.0% on GDPval and 87.3% on SpreadsheetBench, with long-context and tool-search boundaries.
- Evidence
- Vendor report
- Boundary
- “GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.
Thomas Wiegold BlogIndependent measurement
GPT-5.4: A Four-Model Comparison of Atomic Clock Applications
With one one-shot atomic-clock prompt, GPT-5.4 looked best but synchronization drifted; the article calls this a single-task observation.
- Evidence
- Independent measurement
- Boundary
- “GPT-5.4: A Four-Model Comparison of Atomic Clock Applications” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.
Reddit r/AIAgentsPersonal experience
GPT-5.4: Reddit AI Agents — Multi-step Agents and Model Routing Experience
A four-day Reddit discussion reports GPT-5.4 helping with review and multi-step execution, alongside forgotten constraints, excessive calls, and xhigh cost.
- Evidence
- Personal experience
- Boundary
- “GPT-5.4: Reddit AI Agents — Multi-step Agents and Model Routing Experience” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.