OpenAI reports GPT-5.2 professional-work, spreadsheet, and coding results, including 55.6% on SWE-Bench Pro; full harnesses are not public.
OpenAI News / GPT-5.2 release notes · Read evidenceGPT-5.2 Chat · Reviews and evidence
Which GPT-5.2 Chat conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
The SWE-bench leaderboard compares coding agents by submission and task set; rankings change with versions, scaffolds, and evaluation settings.
SWE-bench Leaderboards · Read evidencePost-launch Reddit discussion has mixed GPT-5.2 coding and conversation reports, without a common task set, snapshot, or control group.
Reddit / r/ChatGPT · Read evidenceFull reviews and related reading
Selected evidence
GPT-5.2 Family: Official Release Benchmarks and Chat Positioning
OpenAI reports GPT-5.2 professional-work, spreadsheet, and coding results, including 55.6% on SWE-Bench Pro; full harnesses are not public.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
SWE-bench Leaderboard: Comparing GPT-5.2 Coding Agents
The SWE-bench leaderboard compares coding agents by submission and task set; rankings change with versions, scaffolds, and evaluation settings.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
Reddit Users' Coding and Conversation Experience After the GPT-5.2 Launch
Post-launch Reddit discussion has mixed GPT-5.2 coding and conversation reports, without a common task set, snapshot, or control group.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
All sources
All sources
GPT-5.2 Family: Official Release Benchmarks and Chat Positioning
OpenAI reports GPT-5.2 professional-work, spreadsheet, and coding results, including 55.6% on SWE-Bench Pro; full harnesses are not public.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
SWE-bench Leaderboard: Comparing GPT-5.2 Coding Agents
The SWE-bench leaderboard compares coding agents by submission and task set; rankings change with versions, scaffolds, and evaluation settings.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
Reddit Users' Coding and Conversation Experience After the GPT-5.2 Launch
Post-launch Reddit discussion has mixed GPT-5.2 coding and conversation reports, without a common task set, snapshot, or control group.
Unverified: the original source could not be rechecked.
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
GPT-5.2 Chat
Compare GPT-5.2 Chat in Tabbit
Model access, features, and permissions depend on your current client account.