TabbitBlog

GPT-5.4: What It Is, What Changed, and How to Access It

A sourced GPT-5.4 overview covering native computer use, professional work, tool search, context and billing limits, access routes, and practical risks.

In this article
  1. Key takeaways
  2. GPT-5.4 at a glance
  3. What changed from GPT-5.3-Codex and GPT-5.2?
  4. Access without mixing products
  5. Choose a scenario and run one check
  6. Community evidence: speed, control and route differences
  7. What Tabbit can and cannot establish
  8. Safety and unknown risks
  9. Verdict
  10. Sources and further questions
  11. Is GPT-5.4 the same in ChatGPT, Codex and the API?
  12. Does the API really support 1.05M context?
  13. What reasoning effort should I choose?
  14. Does native computer use make the agent safe?
  15. Does tool search reduce the bill automatically?
  16. Can I use GPT-5.4 in Tabbit?

GPT-5.4 is OpenAI's March 2026 frontier model for professional work, coding, tool use and computer-use agents. It is a good candidate when a workflow must operate across files, applications or websites, but its route and billing details matter as much as the model name.

The decision anchor is easy to miss: the API page lists a 1,050,000-token context, while requests above 272K input tokens are charged at 2x input and 1.5x output for the full session. That is an API billing rule, not a promise that ChatGPT, Codex, a cloud provider or Tabbit exposes the same effective context. Start with the GPT-5.4 model resource, then verify the product and account you will actually use.

Key takeaways

  • GPT-5.4 brings GPT-5.3-Codex coding capabilities into a general-purpose model with professional knowledge work and native computer use.

  • The API lists gpt-5.4, 1,050,000 context, 128,000 maximum output and reasoning effort from none through xhigh.

  • OpenAI reports 75.0% on OSWorld-Verified, 67.3% on WebArena-Verified, 92.8% on Online-Mind2Web and a 47% token reduction from tool search on a 250-task MCP Atlas evaluation.

  • API pricing is $2.50 input, $0.25 cached input and $15 output per million tokens. Over 272K input changes the full-session multiplier; subscription allowances are separate.

  • Native computer use expands what an agent can do, not what it is allowed to do. Confirmation policies, logs and human review remain part of the workflow.

GPT-5.4 at a glance

The GPT-5.4 prompt collection and review collection keep source-level material separate from this overview. They do not grant access to any OpenAI or browser product.

QuestionCurrent API snapshotBoundary
API model IDgpt-5.4Pin the exact ID in logs; ChatGPT and Codex labels are different surfaces.
ReleaseMarch 5, 2026Comparisons should keep model version, date and harness.
Context / maximum output1,050,000 / 128,000 tokensAPI ceiling; effective client windows can be smaller.
Inputs / outputText and images / textTools, computer use and connectors are route-specific.
ReasoningNone, low, medium, high, xhighA UI or plan may expose fewer settings.
API price$2.50 input, $0.25 cached input, $15 output / MTokTool-call fees, Batch/Flex and priority rates are separate.
Long-input boundaryAbove 272K input: 2x input and 1.5x output for the full requestBilling rule, not a universal subscription rule.
Knowledge cutoffAugust 31, 2025 on the API pageCurrent facts need retrieval or verification.

What changed from GPT-5.3-Codex and GPT-5.2?

OpenAI says GPT-5.4 combines GPT-5.3-Codex coding capability with stronger reasoning, tool use and professional work. It is designed to move between spreadsheets, presentations, documents, code and browser tasks without requiring a separate specialist model.

The release describes three concrete changes:

  1. Native computer use: In the API and Codex, GPT-5.4 can operate a computer through screenshots and mouse/keyboard actions, as well as write code for Playwright. Developers can configure confirmation policies for different risk levels.

  2. Tool search: Instead of placing every MCP function definition in the prompt, an agent can search for the needed tool and load its definition when required. OpenAI says this reduced total token use by 47% at the same accuracy on 250 MCP Atlas tasks with 36 MCP servers enabled.

  3. Steerable thinking: GPT-5.4 Thinking can provide a plan before a long answer and accept direction while it works. In Codex, Fast mode is the same model with up to 1.5× faster token velocity; it is a speed route, not a new model.

ChangeOpenAI's published positionWhat to verify locally
Professional workGDPval 83.0% and internal spreadsheet modeling 87.3%Whether your documents, formulas and acceptance rules match the research setup.
CodingSWE-Bench Pro 57.7%, Terminal-Bench 2.0 75.1%Repository, tool harness, effort, tests and review burden.
Computer useOSWorld-Verified 75.0%; WebArena-Verified 67.3%; Online-Mind2Web 92.8%Screenshot fidelity, permissions, confirmation policy and failure recovery.
Tool ecosystemsMCP Atlas token use reduced 47% with tool searchTool definitions, server latency, failures and provider support.
Long contextGraphwalks 256K–1M accuracy 21.4%; MRCR 512K–1M 36.6%A 1M ceiling is not a 1M-context quality guarantee.

OpenAI reports these evaluations with distinct reasoning levels and research environments. For example, GDPval used xhigh for GPT-5.4, while OmniDocBench used none to represent low-cost, low-latency performance. The release itself says production ChatGPT output can differ from research runs. The agentic reasoning guide explains why a benchmark row is not a completed-task guarantee.

Independent evidence adds useful friction. A native-harness Reddit evaluation had 399 total runs but only 14 direct GPT-5.4 runs, mostly Python data-pipeline and Swift UX/reliability work; it found a directional advantage over GPT-5.3-Codex but explicitly kept large error bars (original report). ORCFLO's May 10 cohort evaluates 32 models on quality, cost and speed with four independent judges, while LayerLens reports a separate ten-test Stratix comparison. Neither is a substitute for your own repository test.

Access without mixing products

RouteWhat it isWhat you must confirm
OpenAI APIgpt-5.4 in the Responses API and SDKsOrganization, region, rate limits, tools, reasoning effort and billing tier.
ChatGPTGPT-5.4 Thinking; Pro is a separate routePlan, workspace policy, picker label and ChatGPT context behavior.
CodexGPT-5.4 for coding workflows, with Fast mode and tool configurationDesktop/CLI version, allowance, context setting and repository permissions.
Cloud/provider routeProvider-specific OpenAI deployment or gatewayModel alias, data policy, region, quota and tool support.
Tabbit BrowserA separate browser client if the live picker exposes the modelSigned-in account, visible model label, effective context and reversible test.

The API price is $2.50 per million input tokens, $0.25 for cached input and $15 for output. Batch and Flex are listed at half the standard rate, while Priority processing is twice the standard rate. Tool-specific calls can carry their own fees. Prompts above 272K input receive the full-session 2x/1.5x multiplier. None of this converts into ChatGPT or Codex subscription credits, and none proves Tabbit access.

Choose a scenario and run one check

If your work looks like thisFirst testAcceptance check
Browser or desktop workflowA reversible form or document task with screenshotsCorrect clicks, no unapproved action, complete audit log.
Multi-file codingA small feature with an existing test commandTests pass, diff stays in scope, no invented files or skipped verification.
Spreadsheet/document workA fixed fixture with formulas and expected cellsValues, formulas, formatting and citations match the fixture.
Tool-rich MCP workflowOne task behind tool search and one with a short tool listSame result, fewer prompt tokens, clear tool failures.
Long researchA bounded source packet with an explicit stop conditionCitations open, claims are supported, contradictions are surfaced.

Record model ID, route, effort, context setting, tools, input/output/cache tokens, wall time, retries and human corrections. Compare cost per accepted result, not token price alone. The GPT-5.6 Terra overview and GPT-5.6 Sol overview are newer family comparisons; the Sol 1M guide is not evidence that 5.4 exposes the same long-context behavior.

Community evidence: speed, control and route differences

The community signal is conditional. A Reddit first-impressions thread describes 5.4 as faster than 5.3-Codex and reports a large refactor where 5.4 high executed a plan while 5.3-Codex reviewed it (thread). The same thread includes a warning that the CLI initially compacted around 258K until 1M was explicitly configured. These are client settings, not a contradiction of the API card.

A second community experiment used the same Express JS project and prompt across 16 worktrees, then built a React note-taking feature with an outline panel, shortcuts and preserved behavior. It is a better template for local comparison than a leaderboard, but its code-quality judge was Claude Opus and “correct” summaries remained subjective (method).

The negative signal is about control. One user reports that 5.4 interrupted an explorer subagent and moved on before convergence, concluding that the model needs a strict plan (critique). A YouTube creator's 95% planning score is useful as a creator-defined scenario, not as a universal planning rate. The AI browser guide helps separate a model's ability from the client that grants it pages, files or actions.

What Tabbit can and cannot establish

This draft did not run a signed-in GPT-5.4 task in Tabbit Browser. It therefore makes no claim about picker visibility, effective context, latency, provider routing, subscription allowance, tool permissions or cost. If the model appears in your live selector, begin with a public or reversible task, record the visible model label and effort, and compare the result with a known acceptance check.

The Tabbit Browser overview describes browser-level workflows. It does not grant OpenAI API credits or change ChatGPT/Codex plan limits. Download only as a separate browser decision:

Tabbit Browser

Safety and unknown risks

OpenAI treats GPT-5.4 as High cyber capability under its Preparedness Framework, with monitoring, trusted access controls and asynchronous blocking on some Zero Data Retention surfaces. OpenAI also notes that improving classifiers can create false positives. Native computer use adds an irreversible-action risk: a model can click, type or navigate, but the application must still decide when confirmation is mandatory.

Other unknowns are practical rather than dramatic. A 1M context can still lose useful signal to distraction or compaction. Tool search can reduce prompt tokens but introduces discovery and server-failure paths. The API page's August 31, 2025 knowledge cutoff means current facts need retrieval. Public benchmark values use different efforts, tools and research environments; they should not be merged into one score.

Verdict

GPT-5.4 is a strong first trial for professional document work, coding agents and computer-use workflows. Its most important change is not one leaderboard number: it joins coding, tools and computer operation in one general-purpose route. The practical catch is the boundary between model ceiling and product reality—especially the >272K API multiplier, route-specific context, tool permissions and subscription allowances.

Pin the route and effort, run a reversible fixture, and measure accepted output, token cost, retries, latency and human corrections separately. Keep newer GPT-5.6 or GPT-6 pages in the comparison when lifecycle or capability matters, and treat Tabbit as a browser runtime only after the live picker and task boundary are confirmed.

Sources and further questions

Primary sources are OpenAI's GPT-5.4 announcement, API model page and deployment safety material. Independent sources are LayerLens, ORCFLO and Tom's Guide.

Is GPT-5.4 the same in ChatGPT, Codex and the API?

No. The API ID is gpt-5.4, ChatGPT exposes GPT-5.4 Thinking, and Codex adds its own context, effort, allowance and tool controls. Check the live route rather than assuming the labels are interchangeable.

Does the API really support 1.05M context?

The API model page lists 1,050,000 tokens. That is a model-level ceiling; a client or provider can compact earlier, and the API charges a full-session multiplier above 272K input tokens.

What reasoning effort should I choose?

Start at the lowest effort that passes the acceptance test. Higher effort can improve difficult work but may increase tokens, time and subscription allowance use.

Does native computer use make the agent safe?

No. It expands the actions available to the agent. Use confirmation policies, scoped permissions, logs, tests and human approval for irreversible actions.

Does tool search reduce the bill automatically?

OpenAI reports a 47% token reduction in one 250-task MCP Atlas setup. Your tool definitions, server latency, failures and route can produce a different result.

Can I use GPT-5.4 in Tabbit?

This draft did not verify a signed-in Tabbit session. Check the live selector and treat one successful task as a local observation, not a platform guarantee.

FAQ

What is GPT-5.4?

GPT-5.4 is OpenAI's March 2026 frontier model for professional work, coding, tool use and computer-use agents. It is available as gpt-5.4 in the API, as GPT-5.4 Thinking in ChatGPT, and in Codex.

What are GPT-5.4's context and output limits?

The API model page lists a 1,050,000-token context window and 128,000 maximum output tokens. A client, plan or provider can expose a smaller effective window, and input above 272K has a different full-session price multiplier.

How much does GPT-5.4 cost?

The API lists $2.50 per million input tokens, $0.25 per million cached input tokens and $15 per million output tokens. Prompts above 272K input are charged at 2x input and 1.5x output for the full session.

What changed in GPT-5.4?

OpenAI combines GPT-5.3-Codex coding capabilities with professional knowledge work, native computer use, tool search, improved web search and steerable thinking. The release also adds a preamble and mid-response direction in ChatGPT Thinking.

Where can I use GPT-5.4?

OpenAI lists the API, ChatGPT Thinking and Codex, with additional provider and workspace rules. Model ID, region, organization, plan, app version and tool permissions still determine the route you actually receive.

Can I use GPT-5.4 in Tabbit Browser?

This draft did not run an authenticated Tabbit GPT-5.4 task, so it makes no claim about picker visibility, effective context, latency, cost or tools. Check your live selector and verify one reversible task in your own account.

Take the next step

Let Tabbit work alongside you.

Research across tabs, automate repetitive browser work, and keep every piece of context within reach.