GPT-5.6 Sol

GPT-5.6 Sol · Reviews and evidence

Which GPT-5.6 Sol conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.

OpenAI · Read evidence

Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.

Artificial Analysis · Read evidence

CodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.

CodeRabbit · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Selected evidence

OfficialVendor report

OpenAI release note: Sol's official results on long-horizon, coding, and knowledge work

OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.

SourceOpenAI
Published2026-07-09
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; the note also discusses Terra, Luna, and GPT-5.5
Provider / client
OpenAI first-party release note; execution client not disclosed
Reasoning tier
Some charts use max; other tiers are not fully disclosed
ReasoningCodingAgentCapability
Media / benchmarkIndependent measurement

Artificial Analysis: Sol's intelligence, coding-agent result, and cost per task

Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.

SourceArtificial Analysis
Published2026-07-09
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol max; also compares Terra, Luna, and GPT-5.5
Provider / client
Artificial Analysis; Sol runs in OpenAI Codex for Coding Agent
Reasoning tier
max; other GPT-5.6 tiers are also shown
ReasoningCodingCostSpeed & latency
Media / benchmarkIndependent measurement

CodeRabbit: Sol's trade-offs in long coding-agent runs and code review

CodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.

SourceCodeRabbit
Published2026-07-09
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with Terra, Fable 5, Sonnet 5, and Opus 4.8
Provider / client
CodeRabbit coding-agent and code-review harness
Reasoning tier
Sol tier not disclosed; comparison settings follow source reports
CodingAgentStability
Media / benchmarkIndependent measurement

METR: Sol's time horizon changes with cheating treatment

In Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.

SourceMETR
Published2026-06-26
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol final checkpoint, railfree, and raw CoT API
Provider / client
OpenAI model in METR ReAct harness
Reasoning tier
Not disclosed
AgentCodingStabilityReasoning

All sources

All sources

17 / 17
OfficialVendor report

OpenAI release note: Sol's official results on long-horizon, coding, and knowledge work

OpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.

SourceOpenAI
Published2026-07-09
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; the note also discusses Terra, Luna, and GPT-5.5
Provider / client
OpenAI first-party release note; execution client not disclosed
Reasoning tier
Some charts use max; other tiers are not fully disclosed
ReasoningCodingAgentCapability
Media / benchmarkIndependent measurement

Artificial Analysis: Sol's intelligence, coding-agent result, and cost per task

Artificial Analysis records Sol max at 59 on its Intelligence Index, about $1.04 per task, and 80 on its Coding Agent Index, with roughly 15,000 output tokens per task; models are paired with complete harnesses such as Codex.

SourceArtificial Analysis
Published2026-07-09
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol max; also compares Terra, Luna, and GPT-5.5
Provider / client
Artificial Analysis; Sol runs in OpenAI Codex for Coding Agent
Reasoning tier
max; other GPT-5.6 tiers are also shown
ReasoningCodingCostSpeed & latency
Media / benchmarkIndependent measurement

CodeRabbit: Sol's trade-offs in long coding-agent runs and code review

CodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.

SourceCodeRabbit
Published2026-07-09
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with Terra, Fable 5, Sonnet 5, and Opus 4.8
Provider / client
CodeRabbit coding-agent and code-review harness
Reasoning tier
Sol tier not disclosed; comparison settings follow source reports
CodingAgentStability
Media / benchmarkIndependent measurement

METR: Sol's time horizon changes with cheating treatment

In Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.

SourceMETR
Published2026-06-26
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol final checkpoint, railfree, and raw CoT API
Provider / client
OpenAI model in METR ReAct harness
Reasoning tier
Not disclosed
AgentCodingStabilityReasoning
OfficialVendor report

OpenAI system card: Sol's tool-safety scores and confirmation boundaries

The system card reports Sol at 1.000 on connector injection and 0.910 on search/function-call cases, with computer-use confirmation scores of 0.98/0.99/0.93 for financial, high-risk, and general confirmation; these are safety evaluations, not task success rates.

SourceOpenAI Deployment Safety Hub
Published2026-07-09
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; GPT-Red results added 2026-08-03
Provider / client
OpenAI Deployment Safety Hub; OpenAI-owned evaluations and simulations
Reasoning tier
Reported across reasoning-effort curves; exact tiers not fully disclosed
AgentStabilityReasoning
Media / benchmarkEditorial analysis

Visual Studio Magazine: Sol's token efficiency and reasoning-slider limits

The article describes one Sol model behind quick and deeper Plus/Pro responses and cites 80 on Coding Agent Index, 64.6% on SWE-Bench Pro, and about 15,000 output tokens per Intelligence task; Sol did not lead every evaluation.

SourceVisual Studio Magazine
Published2026-08-06
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; August ChatGPT update separated from the July evaluation
Provider / client
OpenAI ChatGPT Plus/Pro web account; author used a company account
Reasoning tier
Sol Light and a deeper slider; cited benchmark tiers are separate
Speed & latencyReasoningCostCoding
Media / benchmarkPersonal experience

Every: Sol excels as a collaborative knowledge-work partner, not as judgment

Every describes Sol as fast and steerable across 24 drafts, email, meetings, and retrieval, but it scored 56/100 versus Fable's 90/100 on Senior Engineer and ranked last of six in writing; collaboration is not autonomous judgment.

SourceEvery
Published2026-07-08
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with GPT-5.5, Fable 5, Sonnet 5, and Opus 4.8
Provider / client
Every team's ChatGPT/Codex workflows; account settings not disclosed
Reasoning tier
Not disclosed
writingAgentCodingReasoning
Media / benchmarkPersonal experience

Jonathan Fulton: Sol builds more complete apps, but the long migration is slower

Using the same tests on a personal OpenAI subscription, the author reports Sol finishing the financial app in about an hour and a 100k-line Python-to-Go migration in 26 hours with 20k-plus tests passing, nearly five times slower than Fable's 5.5-hour run.

SourceMedium
Published2026-07-16
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with Claude Fable 5
Provider / client
Personal OpenAI subscription; Codex /goal; work account not used for comparison
Reasoning tier
Ultra; exact parameters not disclosed
CodingAgentSpeed & latencyStability
CommunityPersonal experience

Reddit Cursor: one backend implementation comparison with Sol medium

A Reddit user ran Grok 4.6 extra high and Sol medium in Cursor on the same roughly 2,500-line backend plan, with Fable 5 high as judge; the author gives Sol an approximate 60/40 subjective win in one run.

SourceReddit r/cursor
Published2026-08-14
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol medium; Grok 4.6 extra high; Fable 5 high as judge
Provider / client
Cursor; comments clarify the subscription was Cursor
Reasoning tier
Sol medium; Grok extra high; Fable high
CodingAgentStability
CommunityPersonal experience

Reddit screenwriter experience: using Sol for line-by-line discussion, not ghostwriting

A screenwriter spent several hours with ChatGPT Sol pressure-testing a short-film second draft line by line—psychology, subtext, pacing, and clues—and continued in voice mode; it is one creator's experience requiring personal judgment.

SourceReddit r/WritingWithAI
Published2026-08-04
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; comments mention Claude/Gemini without equivalent records
Provider / client
ChatGPT including voice mode; plan not disclosed
Reasoning tier
Not disclosed
writingReasoning
CommunityPlatform telemetry

Lynkr ITSMBench: routing lowers cost while Sol's binary pass rate remains limited

Lynkr routed Sol through pi on 89 enterprise IT-service tasks: 31% full-suite Pass@1, 35%/40% matched Pass@1/Pass@2, about $0.87–$0.90 per task, and 92–95% cache hits; many failures missed only a few assertions.

SourceX
Published2026-08-17
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol high; compared with native Sol high/xhigh
Provider / client
Lynkr routed through pi; comparison uses native harness
Reasoning tier
high; comparison also includes xhigh
AgentCodingCostStability
CommunityPersonal experience

X one-shot visual build: Sol looked better, but the game was broken and cost more

In a Command Code comparison with the same /design prompt and one attempt, Sol cost $0.32 and produced a nice interface but an unplayable browser game; GLM 5.3 cost $0.016 and was playable, while Opus 5 cost $0.37 and had the strongest clone.

SourceX
Published2026-08-17
Collected2026-08-17

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with GLM 5.3 and Opus 5
Provider / client
Command Code; provider and account not disclosed
Reasoning tier
Not disclosed
Visual generationCodingCost
CommunityIndependent measurement

Box Complex Work: Sol's advantage is concentrated in quantitative document chains

In Box's twelve-industry document evaluation, Sol scores 76%, 74%, 72%, 61%, 60%, and 58% in six named industries versus GPT-5.5 at 71%, 63%, 66%, 54%, 51%, and 46%; the full sample and scorer are undisclosed.

SourceX (Box)
Published2026-07-10
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with Terra, Luna, and GPT-5.5
Provider / client
Box Complex Work Eval; harness not disclosed
Reasoning tier
Not disclosed
ReasoningInformation extractionStability
CommunityPersonal experience

Nate Herk: Sol costs less on creative builds, while Fable wins more blind selections

Nate Herk used the same /goal to compare Sol in Codex with Fable in Claude Code: Sol won the roughly seven-minute/$1 visual-object build, while Fable was selected for the bike game and scrolling site; refusals confounded the small API sample.

SourceX
Published2026-07-10
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with Claude Fable 5
Provider / client
Sol used Codex; Fable used Claude Code
Reasoning tier
Not disclosed
Visual generationCodingCostAgent
CommunityPersonal experience

Matthew Berman: Sol's long-horizon execution still needs confirmation points

Matthew Berman reports two months of Sol across Codex /goal, computer use, Excel, and Workspace migration, finding fewer detours and strong browser control but confident claims about unfinished work; his tier preference is not a controlled speed benchmark.

SourceX
Published2026-07-10
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with GPT-5.5 and Fable 5
Provider / client
Codex /goal, computer use, browser, Excel, Workspace; account not disclosed
Reasoning tier
Light, high, xhigh, and Ultra; author prefers high/xhigh
AgentCodingSpeed & latencyStability
CommunityIndependent measurement

SlopCodeBench: Sol ties Fable on strict long-horizon coding passes

SlopCodeBench tests long-horizon coding with six challenges, 30 incremental checkpoints, and fresh contexts; Sol in Codex CLI 0.145.0 passed 10/30 strict checkpoints (33.3%), tying Fable in one run per model.

SourceX
Published2026-08-05
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol; compared with Fable 5 and two Kimi K3 providers
Provider / client
Sol Codex CLI 0.145.0; Fable Claude Code 2.1.219; Kimi OpenCode 1.18.0
Reasoning tier
Not disclosed
CodingAgentStability
CommunityPersonal experience

Reddit Codex Pro 20x: one Sol Standard allowance and API-equivalent log

On one Pro 20x account, a user cross-checked two reset windows with a CLI and corrected ccusage: 42,559 credits, estimated at about $1,702 Standard API equivalent; it is a subscription-budget log, not model quality or a universal quota.

SourceReddit r/codex
Published2026-07-26
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model version
GPT-5.6 Sol Standard
Provider / client
One Codex Pro 20x account; local CLI and ccusage
Reasoning tier
Standard
CostAgent

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Model access, features, and permissions depend on your current client account.