Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.4 · Official source · Vendor report

GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark

OpenAI reports GPT-5.4 results including 83.0% on GDPval and 87.3% on SpreadsheetBench, with long-context and tool-search boundaries.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Condition
Model/version follows the source; reopened 2026-09-20.
Condition
Task, harness, and sample follow the source; undisclosed fields remain unknown.

Key data and applicable tasks

One-sentence takeaway

Official data supports a broad upgrade for GPT-5.4 in computer use, web search, professional knowledge work, and tool calling, but 1M context, long-context pricing, and differences between research and production environments must be included in reproduction and cost assessments.

Use cases

  • Suitable tasks: Cross-software computer use, professional documents/spreadsheets/presentations, deep web search, tool-intensive Agents, coding, and long-running tasks.

  • Unsuitable tasks: Low-latency tasks that do not require reasoning but default to xhigh, or workflows that treat 1M context as lossless memory throughout.

  • Applicable model versions: gpt-5.4; the release page also includes gpt-5.4-pro, but this note focuses on the standard version.

  • Applicable clients, Agents, or APIs: ChatGPT Thinking, API, Codex; API model name gpt-5.4.

  • Recommended reasoning level and parameters: Most official evaluations use xhigh, but production professional tasks should sweep from none/low/medium according to the official guidance; long-context and tool settings must be fixed.

Test environment and input/configuration

  • General evaluation: The official research environment; the release page explicitly says the results may differ somewhat from ChatGPT in production.

  • GDPval: Knowledge work across 44 occupations in 9 high-GDP industries; GPT-5.4 used xhigh, while GPT-5.2 used the lower heavy available in ChatGPT.

  • OSWorld-Verified: A desktop environment with screenshot plus keyboard/mouse operations.

  • WebArena-Verified: Driven by both the DOM and screenshots; Online-Mind2Web uses screenshot-only observation.

  • MCP Atlas: 36 MCP servers, comparing full definitions with tool search, averaged across 250 tasks.

  • Long context: Graphwalks and MRCR v2; the model supports 1M, while requests above 272K have different prices/limits.

Results

BenchmarkGPT-5.4GPT-5.3-CodexGPT-5.2
GDPval (win or tie)83.0%70.9%70.9%
SWE-Bench Pro Public57.7%56.8%55.6%
OSWorld-Verified75.0%74.0%47.3%
Toolathlon54.6%51.9%46.3%
BrowseComp82.7%77.3%65.8%
WebArena-Verified (DOM + screenshots)67.3%Not disclosedGPT-5.2 65.4%
Online-Mind2Web (screenshots)92.8%Not disclosedChatGPT Atlas 70.9%
MMMU-Pro (without tools/with tools)81.2% / 82.1%Not disclosed79.5% / 80.4%
MCP Atlas67.2%Not disclosed60.6%
GPT-5.4 OpenAI MRCR 512K–1M36.6%Not disclosed—

Professional work additions: GDPval 83.0%; internal investment-banking modeling average 87.3%; FinanceAgent v1.1 56.0%; OfficeQA 68.1%. The release page also reports that, compared with GPT-5.2, GPT-5.4 reduced the error rate for individual statements by 33% and the probability that a complete response contains an error by 18% in a de-identified user-fact error test.

Conclusion

GPT-5.4's relative strengths are in “getting things done”: computer use at 75.0%, BrowseComp at 82.7%, tool search, and long-horizon professional work; compared with GPT-5.2, it improves in coding, documents, tools, and vision. For production Agents, results should be validated with the same tools/harness, cost, and completion rate, rather than relying only on the official top-line xhigh scores.

Limitations

  • Official benchmarks were run by OpenAI; inputs, complete tool schemas, random seeds, failure samples, and confidence intervals have not all been disclosed.

  • BrowseComp scores are affected by the search system, timing, blacklists, and the ChatGPT search tool; the official documentation explicitly says API search may differ.

  • Graphwalks/MRCR results for 1M long context show declining performance at the far end of the context range; “supports 1M” does not mean uniformly reliable performance across the full 1M.

  • Partner evaluations and internal tasks are not independent controlled tests; for example, finance, property-tax portals, and legal benchmarks must be understood within their respective harnesses.

Reproduction steps

  1. Fix the gpt-5.4 snapshot, reasoning effort, verbosity, tool definitions, permissions, and context length.

  2. Select coding, document, spreadsheet, computer-use, and web-search tasks according to the business need; record completion/failure, tool calls, tokens, latency, and human takeovers.

  3. Run tool search and full tool definitions A/B on the same tasks to verify whether token savings preserve accuracy.

  4. Calculate actual costs separately for requests below and above 272K, and run retrieval/citation tests from 256K to 1M.

  5. Compare with GPT-5.2/5.3-Codex or the target model using the same harness, and report task types rather than only the overall average.

Original evidence and data

The official release page provides the table above, test definitions, pricing, 47% token savings from tool search, and the original figures for OSWorld/WebArena/Online-Mind2Web. It also states that most evaluations use xhigh and that results from the research environment may differ from those online.

Source excerpt or observation (compliance short quote only)

OpenAI describes GPT-5.4 as “most capable and efficient frontier model for professional work” while also disclosing the remote retrieval and cost boundaries of 1M long context; both should be recorded together.

What this supports

  • Supports the source-specific finding in “GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark”: OpenAI reports GPT-5.4 results including 83.0% on GDPval and 87.3% on SpreadsheetBench, with long-context and tool-search boundaries.

What this does not support

  • “GPT-5.4: OpenAI's Official Professional Work and Agent Benchmark” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

OpenAI Newsroom · OpenAI official · Original publication date 2026-03-05 · Site edit date 2026-09-20

Open original source

GPT-5.4

Compare GPT-5.4 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.4: What It Is, What Changed, and How to Access It

A sourced GPT-5.4 overview covering native computer use, professional work, tool search, context and billing limits, access routes, and practical risks.

Related reviews

GPT-5.4: A Four-Model Comparison of Atomic Clock ApplicationsWith one one-shot atomic-clock prompt, GPT-5.4 looked best but synchronization drifted; the article calls this a single-task observation.GPT-5.4: Reddit AI Agents — Multi-step Agents and Model Routing ExperienceA four-day Reddit discussion reports GPT-5.4 helping with review and multi-step execution, alongside forgotten constraints, excessive calls, and xhigh cost.GPT-5.4: Result Contracts and Verification Loop PromptGPT-5.4 official guidance supports a result contract and verification loop for long tasks; this detail targets one testable code delivery.GPT-5.4: Responses API Tool Search and Phase ConfigurationThe official pages describe Responses tool search and long-context configuration; this detail limits the workflow to an explicit allowlist and phased context.