GPT-5.2 Family: Official Release Benchmarks and Chat Positioning
OpenAI reports GPT-5.2 professional-work, spreadsheet, and coding results, including 55.6% on SWE-Bench Pro; full harnesses are not public.
Evidence
Vendor report
Boundary
“GPT-5.2 Family: Official Release Benchmarks and Chat Positioning” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.
The SWE-bench leaderboard compares coding agents by submission and task set; rankings change with versions, scaffolds, and evaluation settings.
Evidence
Independent measurement
Boundary
“SWE-bench Leaderboard: Comparing GPT-5.2 Coding Agents” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.
Reddit Users' Coding and Conversation Experience After the GPT-5.2 Launch
Post-launch Reddit discussion has mixed GPT-5.2 coding and conversation reports, without a common task set, snapshot, or control group.
Evidence
Personal experience
Boundary
“Reddit Users' Coding and Conversation Experience After the GPT-5.2 Launch” does not publish a common harness, fixed model snapshot, or independent repeats; the finding cannot establish production success beyond its stated task.