Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGPT-5.5

GPT-5.5 Official Benchmarks, Pricing, and Safety Boundaries

Original source

OpenAI

AuthorOpenAI

Source date2026-04-23

Tabbit curation2026-08-19

Read original

One-sentence takeaway

OpenAI's published data positions GPT-5.5 as a high-end model for tool-intensive coding, browsing, knowledge work, and cross-application agents. However, the figures in the table are official evaluation results, so any reproduction must preserve the harness, tools, and model snapshot.

Use cases

  • Suitable tasks: Complex coding and debugging, terminal agents, browsing research, data analysis, cross-application knowledge work, and long-context tasks.

  • Unsuitable tasks: Safety-critical production workflows that select a model solely by official leaderboard numbers; they must be retested against your own data, permissions, and cost constraints.

  • Applicable model version: GPT-5.5 as referred to on the launch page; the API snapshot is gpt-5.5-2026-04-23, and the model page listed approximately 1.05M tokens of context and a maximum output of 128K tokens at the time of collection.

  • Applicable client, agent, or API: Codex, the Responses API, and agents equipped with browsing, terminal, or code tools.

  • Recommended reasoning level and parameters: The launch page does not provide uniform task parameters; start with medium based on the official model guidance, then adjust it using representative evals.

Test/workflow steps

  1. Record the model ID/snapshot, tool versions, system prompt, context length, reasoning level, and token budget.

  2. Choose benchmarks that match the actual task: terminal coding, software operation, browsing research, knowledge work, or long-context tasks.

  3. Run multiple times with identical inputs and tool permissions, recording the success rate, latency, input/output/reasoning tokens, tool errors, and amount of human revision.

  4. Use the official figures only as directional comparisons; do not fill in harnesses, random seeds, or per-question results that the launch page does not disclose.

Test environment

  • Model/version: GPT-5.5; the launch page lists GPT-5.4, Claude Opus 4.7, and Gemini 3.1 Pro among the comparison models.

  • Tools and runtime environment: Each benchmark uses its own official or internal harness; the launch page does not disclose a complete environment configuration for every project.

  • Inputs and evaluation method: Official benchmark sets and internal tasks; the specific question inputs have not all been disclosed.

Input/configuration

The API information explicitly stated on the official page: GPT-5.5 input pricing is $5 per million tokens and output pricing is $30 per million tokens; the context window is 1M tokens, and Codex context is 400K tokens. API inputs longer than 272K tokens trigger the long-context pricing rules listed on the page; check the current pricing page for the exact terms.

Results data

TestGPT-5.5Comparison/notes
Terminal-Bench 2.082.7GPT-5.4 75.1; Opus 4.7 69.4; Gemini 3.1 68.5
Expert-SWE73.1GPT-5.4 68.5
GDPval84.9GPT-5.4 83.0; Opus 4.7 80.3; Gemini 3.1 67.3
OSWorld78.7GPT-5.4 75.0; Opus 4.7 78.0
Toolathlon55.6The launch page gives a single value; the full configuration is not disclosed
BrowseComp84.4GPT-5.5 Pro 90.1
FrontierMath T1–351.7T4 is also listed at 35.4
CyberGym81.8Only the officially published score is recorded; harmful operational content is not repeated
MRCR 512K–1M74.0Long-context retrieval retention metric

Conclusion

The official results show that GPT-5.5 is highly competitive on terminal coding, software operation, knowledge work, and browsing tasks. This supports putting it on the candidate list for complex agents and cross-tool workflows. It does not prove that the model leads equally on every codebase, browsing task, or enterprise process, and benchmark scores must not be directly converted into business ROI.

Limitations and reproduction steps

  • Limitations: The results were published by the model provider; the tools, prompts, time limits, and graders differ across projects, and some harnesses have not been made public.

  • Reproduction steps: Fix gpt-5.5-2026-04-23 (if it is still available), record the Responses/Codex versions and tool permissions, select comparable public tasks, preserve the original inputs, tool traces, outputs, tokens, and human scores, then run GPT-5.4 or other candidates in the same harness for comparison.

  • Safety boundaries: OpenAI uses stronger classifiers and trusted access controls for biological and cybersecurity capabilities; when reproducing safety evaluations, perform only compliant assessments and do not copy harmful instructions.

Original evidence and data

The launch page's “GPT-5.5 benchmarks” table provides the scores and comparison models above; the pricing, context, and snapshot information is supplemented by OpenAI's model page and API pricing page. This article preserves the division of responsibility between sources and does not present test settings undisclosed by the launch page as facts.

Applicability boundaries

  • The scores are officially reported values and must not be labeled as independently reproduced.

  • When comparing models, prioritize matching the task type, tools, and time budget rather than ranking by a single overall leaderboard.

  • Pricing and snapshots may change; recheck the official model and pricing pages before production deployment.

Source excerpt or observation (compliant short quote only)

The launch page describes GPT-5.5 as “a major step forward in coding, knowledge work, and computer use,” but that positioning should still be validated through evaluation on your own tasks.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-5.5

Use and compare models in Tabbit

GPT-5.5

Related reviews

MediaVellum

GPT-5.5 Vellum Cross-Model Benchmarking and Vendor Data Boundaries

GPT-5.5

Related prompts

OfficialOpenAI Developers

GPT-5.5 Outcome-Oriented and Verification-Driven Agent Prompt

OfficialOpenAI Developers

GPT-5.5 Responses API Reasoning and Multi-turn State Configuration

OfficialOpenAI Developers

GPT-5.5 Official Model Page: API and Tool Configuration