GPT-5.4 is OpenAI's March 2026 frontier model for professional work, coding, tool use and computer-use agents. It is a good candidate when a workflow must operate across files, applications or websites, but its route and billing details matter as much as the model name.
The decision anchor is easy to miss: the API page lists a 1,050,000-token context, while requests above 272K input tokens are charged at 2x input and 1.5x output for the full session. That is an API billing rule, not a promise that ChatGPT, Codex, a cloud provider or Tabbit exposes the same effective context. Start with the GPT-5.4 model resource, then verify the product and account you will actually use.
Key takeaways
GPT-5.4 brings GPT-5.3-Codex coding capabilities into a general-purpose model with professional knowledge work and native computer use.
The API lists
gpt-5.4, 1,050,000 context, 128,000 maximum output and reasoning effort from none through xhigh.OpenAI reports 75.0% on OSWorld-Verified, 67.3% on WebArena-Verified, 92.8% on Online-Mind2Web and a 47% token reduction from tool search on a 250-task MCP Atlas evaluation.
API pricing is $2.50 input, $0.25 cached input and $15 output per million tokens. Over 272K input changes the full-session multiplier; subscription allowances are separate.
Native computer use expands what an agent can do, not what it is allowed to do. Confirmation policies, logs and human review remain part of the workflow.
GPT-5.4 at a glance
The GPT-5.4 prompt collection and review collection keep source-level material separate from this overview. They do not grant access to any OpenAI or browser product.
| Question | Current API snapshot | Boundary |
|---|---|---|
| API model ID | gpt-5.4 | Pin the exact ID in logs; ChatGPT and Codex labels are different surfaces. |
| Release | March 5, 2026 | Comparisons should keep model version, date and harness. |
| Context / maximum output | 1,050,000 / 128,000 tokens | API ceiling; effective client windows can be smaller. |
| Inputs / output | Text and images / text | Tools, computer use and connectors are route-specific. |
| Reasoning | None, low, medium, high, xhigh | A UI or plan may expose fewer settings. |
| API price | $2.50 input, $0.25 cached input, $15 output / MTok | Tool-call fees, Batch/Flex and priority rates are separate. |
| Long-input boundary | Above 272K input: 2x input and 1.5x output for the full request | Billing rule, not a universal subscription rule. |
| Knowledge cutoff | August 31, 2025 on the API page | Current facts need retrieval or verification. |
What changed from GPT-5.3-Codex and GPT-5.2?
OpenAI says GPT-5.4 combines GPT-5.3-Codex coding capability with stronger reasoning, tool use and professional work. It is designed to move between spreadsheets, presentations, documents, code and browser tasks without requiring a separate specialist model.
The release describes three concrete changes:
Native computer use: In the API and Codex, GPT-5.4 can operate a computer through screenshots and mouse/keyboard actions, as well as write code for Playwright. Developers can configure confirmation policies for different risk levels.
Tool search: Instead of placing every MCP function definition in the prompt, an agent can search for the needed tool and load its definition when required. OpenAI says this reduced total token use by 47% at the same accuracy on 250 MCP Atlas tasks with 36 MCP servers enabled.
Steerable thinking: GPT-5.4 Thinking can provide a plan before a long answer and accept direction while it works. In Codex, Fast mode is the same model with up to 1.5× faster token velocity; it is a speed route, not a new model.
| Change | OpenAI's published position | What to verify locally |
|---|---|---|
| Professional work | GDPval 83.0% and internal spreadsheet modeling 87.3% | Whether your documents, formulas and acceptance rules match the research setup. |
| Coding | SWE-Bench Pro 57.7%, Terminal-Bench 2.0 75.1% | Repository, tool harness, effort, tests and review burden. |
| Computer use | OSWorld-Verified 75.0%; WebArena-Verified 67.3%; Online-Mind2Web 92.8% | Screenshot fidelity, permissions, confirmation policy and failure recovery. |
| Tool ecosystems | MCP Atlas token use reduced 47% with tool search | Tool definitions, server latency, failures and provider support. |
| Long context | Graphwalks 256K–1M accuracy 21.4%; MRCR 512K–1M 36.6% | A 1M ceiling is not a 1M-context quality guarantee. |
OpenAI reports these evaluations with distinct reasoning levels and research environments. For example, GDPval used xhigh for GPT-5.4, while OmniDocBench used none to represent low-cost, low-latency performance. The release itself says production ChatGPT output can differ from research runs. The agentic reasoning guide explains why a benchmark row is not a completed-task guarantee.
Independent evidence adds useful friction. A native-harness Reddit evaluation had 399 total runs but only 14 direct GPT-5.4 runs, mostly Python data-pipeline and Swift UX/reliability work; it found a directional advantage over GPT-5.3-Codex but explicitly kept large error bars (original report). ORCFLO's May 10 cohort evaluates 32 models on quality, cost and speed with four independent judges, while LayerLens reports a separate ten-test Stratix comparison. Neither is a substitute for your own repository test.
Access without mixing products
| Route | What it is | What you must confirm |
|---|---|---|
| OpenAI API | gpt-5.4 in the Responses API and SDKs | Organization, region, rate limits, tools, reasoning effort and billing tier. |
| ChatGPT | GPT-5.4 Thinking; Pro is a separate route | Plan, workspace policy, picker label and ChatGPT context behavior. |
| Codex | GPT-5.4 for coding workflows, with Fast mode and tool configuration | Desktop/CLI version, allowance, context setting and repository permissions. |
| Cloud/provider route | Provider-specific OpenAI deployment or gateway | Model alias, data policy, region, quota and tool support. |
| Tabbit Browser | A separate browser client if the live picker exposes the model | Signed-in account, visible model label, effective context and reversible test. |
The API price is $2.50 per million input tokens, $0.25 for cached input and $15 for output. Batch and Flex are listed at half the standard rate, while Priority processing is twice the standard rate. Tool-specific calls can carry their own fees. Prompts above 272K input receive the full-session 2x/1.5x multiplier. None of this converts into ChatGPT or Codex subscription credits, and none proves Tabbit access.
Choose a scenario and run one check
| If your work looks like this | First test | Acceptance check |
|---|---|---|
| Browser or desktop workflow | A reversible form or document task with screenshots | Correct clicks, no unapproved action, complete audit log. |
| Multi-file coding | A small feature with an existing test command | Tests pass, diff stays in scope, no invented files or skipped verification. |
| Spreadsheet/document work | A fixed fixture with formulas and expected cells | Values, formulas, formatting and citations match the fixture. |
| Tool-rich MCP workflow | One task behind tool search and one with a short tool list | Same result, fewer prompt tokens, clear tool failures. |
| Long research | A bounded source packet with an explicit stop condition | Citations open, claims are supported, contradictions are surfaced. |
Record model ID, route, effort, context setting, tools, input/output/cache tokens, wall time, retries and human corrections. Compare cost per accepted result, not token price alone. The GPT-5.6 Terra overview and GPT-5.6 Sol overview are newer family comparisons; the Sol 1M guide is not evidence that 5.4 exposes the same long-context behavior.
Community evidence: speed, control and route differences
The community signal is conditional. A Reddit first-impressions thread describes 5.4 as faster than 5.3-Codex and reports a large refactor where 5.4 high executed a plan while 5.3-Codex reviewed it (thread). The same thread includes a warning that the CLI initially compacted around 258K until 1M was explicitly configured. These are client settings, not a contradiction of the API card.
A second community experiment used the same Express JS project and prompt across 16 worktrees, then built a React note-taking feature with an outline panel, shortcuts and preserved behavior. It is a better template for local comparison than a leaderboard, but its code-quality judge was Claude Opus and “correct” summaries remained subjective (method).
The negative signal is about control. One user reports that 5.4 interrupted an explorer subagent and moved on before convergence, concluding that the model needs a strict plan (critique). A YouTube creator's 95% planning score is useful as a creator-defined scenario, not as a universal planning rate. The AI browser guide helps separate a model's ability from the client that grants it pages, files or actions.
What Tabbit can and cannot establish
This draft did not run a signed-in GPT-5.4 task in Tabbit Browser. It therefore makes no claim about picker visibility, effective context, latency, provider routing, subscription allowance, tool permissions or cost. If the model appears in your live selector, begin with a public or reversible task, record the visible model label and effort, and compare the result with a known acceptance check.
The Tabbit Browser overview describes browser-level workflows. It does not grant OpenAI API credits or change ChatGPT/Codex plan limits. Download only as a separate browser decision:
Safety and unknown risks
OpenAI treats GPT-5.4 as High cyber capability under its Preparedness Framework, with monitoring, trusted access controls and asynchronous blocking on some Zero Data Retention surfaces. OpenAI also notes that improving classifiers can create false positives. Native computer use adds an irreversible-action risk: a model can click, type or navigate, but the application must still decide when confirmation is mandatory.
Other unknowns are practical rather than dramatic. A 1M context can still lose useful signal to distraction or compaction. Tool search can reduce prompt tokens but introduces discovery and server-failure paths. The API page's August 31, 2025 knowledge cutoff means current facts need retrieval. Public benchmark values use different efforts, tools and research environments; they should not be merged into one score.
Verdict
GPT-5.4 is a strong first trial for professional document work, coding agents and computer-use workflows. Its most important change is not one leaderboard number: it joins coding, tools and computer operation in one general-purpose route. The practical catch is the boundary between model ceiling and product reality—especially the >272K API multiplier, route-specific context, tool permissions and subscription allowances.
Pin the route and effort, run a reversible fixture, and measure accepted output, token cost, retries, latency and human corrections separately. Keep newer GPT-5.6 or GPT-6 pages in the comparison when lifecycle or capability matters, and treat Tabbit as a browser runtime only after the live picker and task boundary are confirmed.
Sources and further questions
Primary sources are OpenAI's GPT-5.4 announcement, API model page and deployment safety material. Independent sources are LayerLens, ORCFLO and Tom's Guide.
Is GPT-5.4 the same in ChatGPT, Codex and the API?
No. The API ID is gpt-5.4, ChatGPT exposes GPT-5.4 Thinking, and Codex adds its own context, effort, allowance and tool controls. Check the live route rather than assuming the labels are interchangeable.
Does the API really support 1.05M context?
The API model page lists 1,050,000 tokens. That is a model-level ceiling; a client or provider can compact earlier, and the API charges a full-session multiplier above 272K input tokens.
What reasoning effort should I choose?
Start at the lowest effort that passes the acceptance test. Higher effort can improve difficult work but may increase tokens, time and subscription allowance use.
Does native computer use make the agent safe?
No. It expands the actions available to the agent. Use confirmation policies, scoped permissions, logs, tests and human approval for irreversible actions.
Does tool search reduce the bill automatically?
OpenAI reports a 47% token reduction in one 250-task MCP Atlas setup. Your tool definitions, server latency, failures and route can produce a different result.
Can I use GPT-5.4 in Tabbit?
This draft did not verify a signed-in Tabbit session. Check the live selector and treat one successful task as a local observation, not a platform guarantee.
FAQ
What is GPT-5.4?
GPT-5.4 is OpenAI's March 2026 frontier model for professional work, coding, tool use and computer-use agents. It is available as gpt-5.4 in the API, as GPT-5.4 Thinking in ChatGPT, and in Codex.
What are GPT-5.4's context and output limits?
The API model page lists a 1,050,000-token context window and 128,000 maximum output tokens. A client, plan or provider can expose a smaller effective window, and input above 272K has a different full-session price multiplier.
How much does GPT-5.4 cost?
The API lists $2.50 per million input tokens, $0.25 per million cached input tokens and $15 per million output tokens. Prompts above 272K input are charged at 2x input and 1.5x output for the full session.
What changed in GPT-5.4?
OpenAI combines GPT-5.3-Codex coding capabilities with professional knowledge work, native computer use, tool search, improved web search and steerable thinking. The release also adds a preamble and mid-response direction in ChatGPT Thinking.
Where can I use GPT-5.4?
OpenAI lists the API, ChatGPT Thinking and Codex, with additional provider and workspace rules. Model ID, region, organization, plan, app version and tool permissions still determine the route you actually receive.
Can I use GPT-5.4 in Tabbit Browser?
This draft did not run an authenticated Tabbit GPT-5.4 task, so it makes no claim about picker visibility, effective context, latency, cost or tools. Check your live selector and verify one reversible task in your own account.