GPT-5.5

GPT-5.5 · Reviews and evidence

Which GPT-5.5 conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

OpenAI's release page attributes GPT-5.5 results in tool-heavy coding, browsing, and cross-software agents to specific harnesses; the tables cannot be reproduced without the snapshot and tools.

OpenAI · Read evidence

Vellum aggregates public GPT-5.5, Claude, and Gemini scores and shows substantial task-to-task variation; it is a cross-source digest, not a controlled rerun.

Vellum · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

GPT-5.5: What It Is, Access, Pricing, and the Migration Deadline

A sourced GPT-5.5 overview covering API limits, benchmark boundaries, access routes, pricing, and the announced October 14, 2026 product migration.

Selected evidence

OfficialVendor report

GPT-5.5 Official Benchmarks, Pricing, and Safety Boundaries

OpenAI's release page attributes GPT-5.5 results in tool-heavy coding, browsing, and cross-software agents to specific harnesses; the tables cannot be reproduced without the snapshot and tools.

SourceOpenAI
Published2026-04-23
Collected2026-08-20
Source-specific observation
The 2026-04-23 release uses GPT-5.5 and snapshot gpt-5.5-2026-04-23 with browsing, terminal, coding, and agent tools.
Published conditions
The page lists about 1.05M context and 128K maximum output, but does not publish identical inputs and repeats for each business task.
CapabilityAgent
Media / benchmarkIndependent measurement

GPT-5.5 Vellum Cross-Model Benchmarking and Vendor Data Boundaries

Vellum aggregates public GPT-5.5, Claude, and Gemini scores and shows substantial task-to-task variation; it is a cross-source digest, not a controlled rerun.

SourceVellum
PublishedUnknown
Collected2026-08-20

Unverified: the original source could not be rechecked.

Source-specific observation
The article compares public Terminal, GDPval, OSWorld, SWE Pro, and related results for GPT-5.5, Claude, and Gemini, with source-specific snapshots.
Published conditions
Vellum does not provide one uniform API harness, random seed, and raw output for every task; collected 2026-08-18 and not reopened in this pass.
CapabilityReasoning

All sources

All sources

2 / 2
OfficialVendor report

GPT-5.5 Official Benchmarks, Pricing, and Safety Boundaries

OpenAI's release page attributes GPT-5.5 results in tool-heavy coding, browsing, and cross-software agents to specific harnesses; the tables cannot be reproduced without the snapshot and tools.

SourceOpenAI
Published2026-04-23
Collected2026-08-20
Source-specific observation
The 2026-04-23 release uses GPT-5.5 and snapshot gpt-5.5-2026-04-23 with browsing, terminal, coding, and agent tools.
Published conditions
The page lists about 1.05M context and 128K maximum output, but does not publish identical inputs and repeats for each business task.
CapabilityAgent
Media / benchmarkIndependent measurement

GPT-5.5 Vellum Cross-Model Benchmarking and Vendor Data Boundaries

Vellum aggregates public GPT-5.5, Claude, and Gemini scores and shows substantial task-to-task variation; it is a cross-source digest, not a controlled rerun.

SourceVellum
PublishedUnknown
Collected2026-08-20

Unverified: the original source could not be rechecked.

Source-specific observation
The article compares public Terminal, GDPval, OSWorld, SWE Pro, and related results for GPT-5.5, Claude, and Gemini, with source-specific snapshots.
Published conditions
Vellum does not provide one uniform API harness, random seed, and raw output for every task; collected 2026-08-18 and not reopened in this pass.
CapabilityReasoning

GPT-5.5

Compare GPT-5.5 in Tabbit

Model access, features, and permissions depend on your current client account.