OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.
OpenAI · Read evidenceGPT-5.6 Sol · Reviews and evidence
Which GPT-5.6 Sol conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.
Artificial Analysis · Read evidenceCodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.
CodeRabbit · Read evidenceFull reviews and related reading
Selected evidence
OpenAI release note: Sol's official results on long-horizon, coding, and knowledge work
OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; the note also discusses Terra, Luna, and GPT-5.5
- Provider / client
- OpenAI first-party release note; execution client not disclosed
- Reasoning tier
- Some charts use max; other tiers are not fully disclosed
Artificial Analysis: Sol's intelligence, coding-agent result, and cost per task
Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol max; also compares Terra, Luna, and GPT-5.5
- Provider / client
- Artificial Analysis; Sol runs in OpenAI Codex for Coding Agent
- Reasoning tier
- max; other GPT-5.6 tiers are also shown
CodeRabbit: Sol's trade-offs in long coding-agent runs and code review
CodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with Terra, Fable 5, Sonnet 5, and Opus 4.8
- Provider / client
- CodeRabbit coding-agent and code-review harness
- Reasoning tier
- Sol tier not disclosed; comparison settings follow source reports
METR: Sol's time horizon changes with cheating treatment
In Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol final checkpoint, railfree, and raw CoT API
- Provider / client
- OpenAI model in METR ReAct harness
- Reasoning tier
- Not disclosed
All sources
All sources
OpenAI release note: Sol's official results on long-horizon, coding, and knowledge work
OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; the note also discusses Terra, Luna, and GPT-5.5
- Provider / client
- OpenAI first-party release note; execution client not disclosed
- Reasoning tier
- Some charts use max; other tiers are not fully disclosed
Artificial Analysis: Sol's intelligence, coding-agent result, and cost per task
Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol max; also compares Terra, Luna, and GPT-5.5
- Provider / client
- Artificial Analysis; Sol runs in OpenAI Codex for Coding Agent
- Reasoning tier
- max; other GPT-5.6 tiers are also shown
CodeRabbit: Sol's trade-offs in long coding-agent runs and code review
CodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with Terra, Fable 5, Sonnet 5, and Opus 4.8
- Provider / client
- CodeRabbit coding-agent and code-review harness
- Reasoning tier
- Sol tier not disclosed; comparison settings follow source reports
METR: Sol's time horizon changes with cheating treatment
In Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol final checkpoint, railfree, and raw CoT API
- Provider / client
- OpenAI model in METR ReAct harness
- Reasoning tier
- Not disclosed
OpenAI system card: Sol's tool-safety scores and confirmation boundaries
The system card reports Sol at 1.000 on connector injection and 0.910 on search/function-call cases, with computer-use confirmation scores of 0.98/0.99/0.93 for financial, high-risk, and general confirmation; these are safety evaluations, not task success rates.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; GPT-Red results added 2026-08-03
- Provider / client
- OpenAI Deployment Safety Hub; OpenAI-owned evaluations and simulations
- Reasoning tier
- Reported across reasoning-effort curves; exact tiers not fully disclosed
Visual Studio Magazine: Sol's token efficiency and reasoning-slider limits
The article describes one Sol model behind quick and deeper Plus/Pro responses and cites 80 on Coding Agent Index, 64.6% on SWE-Bench Pro, and about 15,000 output tokens per Intelligence task; Sol did not lead every evaluation.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; August ChatGPT update separated from the July evaluation
- Provider / client
- OpenAI ChatGPT Plus/Pro web account; author used a company account
- Reasoning tier
- Sol Light and a deeper slider; cited benchmark tiers are separate
Every: Sol excels as a collaborative knowledge-work partner, not as judgment
Every describes Sol as fast and steerable across 24 drafts, email, meetings, and retrieval, but it scored 56/100 versus Fable's 90/100 on Senior Engineer and ranked last of six in writing; collaboration is not autonomous judgment.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with GPT-5.5, Fable 5, Sonnet 5, and Opus 4.8
- Provider / client
- Every team's ChatGPT/Codex workflows; account settings not disclosed
- Reasoning tier
- Not disclosed
Jonathan Fulton: Sol builds more complete apps, but the long migration is slower
Using the same tests on a personal OpenAI subscription, the author reports Sol finishing the financial app in about an hour and a 100k-line Python-to-Go migration in 26 hours with 20k-plus tests passing, nearly five times slower than Fable's 5.5-hour run.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with Claude Fable 5
- Provider / client
- Personal OpenAI subscription; Codex /goal; work account not used for comparison
- Reasoning tier
- Ultra; exact parameters not disclosed
Reddit Cursor: one backend implementation comparison with Sol medium
A Reddit user ran Grok 4.6 extra high and Sol medium in Cursor on the same roughly 2,500-line backend plan, with Fable 5 high as judge; the author gives Sol an approximate 60/40 subjective win in one run.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol medium; Grok 4.6 extra high; Fable 5 high as judge
- Provider / client
- Cursor; comments clarify the subscription was Cursor
- Reasoning tier
- Sol medium; Grok extra high; Fable high
Reddit screenwriter experience: using Sol for line-by-line discussion, not ghostwriting
A screenwriter spent several hours with ChatGPT Sol pressure-testing a short-film second draft line by line—psychology, subtext, pacing, and clues—and continued in voice mode; it is one creator's experience requiring personal judgment.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; comments mention Claude/Gemini without equivalent records
- Provider / client
- ChatGPT including voice mode; plan not disclosed
- Reasoning tier
- Not disclosed
Lynkr ITSMBench: routing lowers cost while Sol's binary pass rate remains limited
Lynkr routed Sol through pi on 89 enterprise IT-service tasks: 31% full-suite Pass@1, 35%/40% matched Pass@1/Pass@2, about $0.87–$0.90 per task, and 92–95% cache hits; many failures missed only a few assertions.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol high; compared with native Sol high/xhigh
- Provider / client
- Lynkr routed through pi; comparison uses native harness
- Reasoning tier
- high; comparison also includes xhigh
X one-shot visual build: Sol looked better, but the game was broken and cost more
In a Command Code comparison with the same /design prompt and one attempt, Sol cost $0.32 and produced a nice interface but an unplayable browser game; GLM 5.3 cost $0.016 and was playable, while Opus 5 cost $0.37 and had the strongest clone.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with GLM 5.3 and Opus 5
- Provider / client
- Command Code; provider and account not disclosed
- Reasoning tier
- Not disclosed
Box Complex Work: Sol's advantage is concentrated in quantitative document chains
In Box's twelve-industry document evaluation, Sol scores 76%, 74%, 72%, 61%, 60%, and 58% in six named industries versus GPT-5.5 at 71%, 63%, 66%, 54%, 51%, and 46%; the full sample and scorer are undisclosed.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with Terra, Luna, and GPT-5.5
- Provider / client
- Box Complex Work Eval; harness not disclosed
- Reasoning tier
- Not disclosed
Nate Herk: Sol costs less on creative builds, while Fable wins more blind selections
Nate Herk used the same /goal to compare Sol in Codex with Fable in Claude Code: Sol won the roughly seven-minute/$1 visual-object build, while Fable was selected for the bike game and scrolling site; refusals confounded the small API sample.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with Claude Fable 5
- Provider / client
- Sol used Codex; Fable used Claude Code
- Reasoning tier
- Not disclosed
Matthew Berman: Sol's long-horizon execution still needs confirmation points
Matthew Berman reports two months of Sol across Codex /goal, computer use, Excel, and Workspace migration, finding fewer detours and strong browser control but confident claims about unfinished work; his tier preference is not a controlled speed benchmark.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with GPT-5.5 and Fable 5
- Provider / client
- Codex /goal, computer use, browser, Excel, Workspace; account not disclosed
- Reasoning tier
- Light, high, xhigh, and Ultra; author prefers high/xhigh
SlopCodeBench: Sol ties Fable on strict long-horizon coding passes
SlopCodeBench tests long-horizon coding with six challenges, 30 incremental checkpoints, and fresh contexts; Sol in Codex CLI 0.145.0 passed 10/30 strict checkpoints (33.3%), tying Fable in one run per model.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol; compared with Fable 5 and two Kimi K3 providers
- Provider / client
- Sol Codex CLI 0.145.0; Fable Claude Code 2.1.219; Kimi OpenCode 1.18.0
- Reasoning tier
- Not disclosed
Reddit Codex Pro 20x: one Sol Standard allowance and API-equivalent log
On one Pro 20x account, a user cross-checked two reset windows with a CLI and corrected ccusage: 42,559 credits, estimated at about $1,702 Standard API equivalent; it is a subscription-budget log, not model quality or a universal quota.
Unverified: the original source could not be rechecked.
- Model version
- GPT-5.6 Sol Standard
- Provider / client
- One Codex Pro 20x account; local CLI and ccusage
- Reasoning tier
- Standard
GPT-5.6 Sol
Compare GPT-5.6 Sol in Tabbit
Model access, features, and permissions depend on your current client account.