OpenAI's release page attributes GPT-5.5 results in tool-heavy coding, browsing, and cross-software agents to specific harnesses; the tables cannot be reproduced without the snapshot and tools.
OpenAI · Read evidenceGPT-5.5 · Reviews and evidence
Which GPT-5.5 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
Vellum aggregates public GPT-5.5, Claude, and Gemini scores and shows substantial task-to-task variation; it is a cross-source digest, not a controlled rerun.
Vellum · Read evidenceFull reviews and related reading
Selected evidence
GPT-5.5 Official Benchmarks, Pricing, and Safety Boundaries
OpenAI's release page attributes GPT-5.5 results in tool-heavy coding, browsing, and cross-software agents to specific harnesses; the tables cannot be reproduced without the snapshot and tools.
- Source-specific observation
- The 2026-04-23 release uses GPT-5.5 and snapshot gpt-5.5-2026-04-23 with browsing, terminal, coding, and agent tools.
- Published conditions
- The page lists about 1.05M context and 128K maximum output, but does not publish identical inputs and repeats for each business task.
GPT-5.5 Vellum Cross-Model Benchmarking and Vendor Data Boundaries
Vellum aggregates public GPT-5.5, Claude, and Gemini scores and shows substantial task-to-task variation; it is a cross-source digest, not a controlled rerun.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The article compares public Terminal, GDPval, OSWorld, SWE Pro, and related results for GPT-5.5, Claude, and Gemini, with source-specific snapshots.
- Published conditions
- Vellum does not provide one uniform API harness, random seed, and raw output for every task; collected 2026-08-18 and not reopened in this pass.
All sources
All sources
GPT-5.5 Official Benchmarks, Pricing, and Safety Boundaries
OpenAI's release page attributes GPT-5.5 results in tool-heavy coding, browsing, and cross-software agents to specific harnesses; the tables cannot be reproduced without the snapshot and tools.
- Source-specific observation
- The 2026-04-23 release uses GPT-5.5 and snapshot gpt-5.5-2026-04-23 with browsing, terminal, coding, and agent tools.
- Published conditions
- The page lists about 1.05M context and 128K maximum output, but does not publish identical inputs and repeats for each business task.
GPT-5.5 Vellum Cross-Model Benchmarking and Vendor Data Boundaries
Vellum aggregates public GPT-5.5, Claude, and Gemini scores and shows substantial task-to-task variation; it is a cross-source digest, not a controlled rerun.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The article compares public Terminal, GDPval, OSWorld, SWE Pro, and related results for GPT-5.5, Claude, and Gemini, with source-specific snapshots.
- Published conditions
- Vellum does not provide one uniform API harness, random seed, and raw output for every task; collected 2026-08-18 and not reopened in this pass.
GPT-5.5
Compare GPT-5.5 in Tabbit
Model access, features, and permissions depend on your current client account.