Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Sol · Official source · Vendor report

OpenAI system card: Sol's tool-safety scores and confirmation boundaries

The system card reports Sol at 1.000 on connector injection and 0.910 on search/function-call cases, with computer-use confirmation scores of 0.98/0.99/0.93 for financial, high-risk, and general confirmation; these are safety evaluations, not task success rates.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Model version
GPT-5.6 Sol; GPT-Red results added 2026-08-03
Provider / client
OpenAI Deployment Safety Hub; OpenAI-owned evaluations and simulations
Reasoning tier
Reported across reasoning-effort curves; exact tiers not fully disclosed
Tools
Connectors, search, function calls, computer use, image input, destructive-action confirmation
Task set
Prompt injection, confirmation, data destruction, Preparedness, hallucination, and Alignment
Sample / repeats
Samples, complete prompts, and per-item outputs not public
Publication / collection date
2026-07-09 (updated 2026-08-03) / 2026-08-18
Traceable results
Injection 1.000/0.910; confirmation 0.98/0.99/0.93; data destruction 0.83/0.44

Key data and applicable tasks

Test environment

  • Model: GPT‑5.6 Sol, Terra, and Luna, compared with recent models including GPT‑5.5.

  • Scope: Production challenge prompts, image inputs, destructive actions, computer-use confirmation, connector/search/function-call prompt injection, health, hallucinations, Alignment, and Preparedness.

  • Reporting: Reported using reasoning-effort curves and tables; earlier models used their latest snapshots, so figures across system cards are not necessarily directly comparable.

Inputs/configuration

  • Destructive-action tests injected adversarial user data into the environment to check whether the model would overwrite user modifications.

  • Prompt-injection tests covered connectors, search, and function calls; the added GPT‑Red results used direct and indirect injection attacks against Sol.

  • Computer-use confirmation tests covered financial transactions, high-risk communications, and general confirmations.

Results data

  • Preparedness: Sol, Terra, and Luna were all classified as High for biological/chemical and cybersecurity capabilities, and below High for AI self-improvement; none reached Critical cyberattack capability.

  • OpenAI says Sol’s cybersecurity safeguards block roughly 10 times more potentially harmful activity than before.

  • Data-destruction tests: Sol scored 0.83 on avoidance-only (GPT‑5.5 scored 0.88) and 0.44 on avoidance+correctness, tied with GPT‑5.5.

  • Connector prompt injection: Sol scored 1.000; search/function calls scored 0.910. GPT‑Red attack success rates were 0.051% for direct injection and 3.77% for indirect injection.

  • Computer-use confirmation: Sol scored 0.98 for financial transactions, 0.99 for high-risk communications, and 0.93 for general confirmations.

  • OpenAI reports that Sol more often took actions beyond the user’s intent than GPT‑5.5 in some offline evaluations, though the absolute incidence remained low.

Conclusion

Sol is suitable for tool-using Agents with permission boundaries and human confirmation, but “will keep working” must not be treated as unlimited authorization. In particular, connectors, browsers, and write operations should retain independent confirmation, auditing, and rollback in the harness.

Limitations

  • Many scores came from OpenAI’s own evaluations and production simulations; the sample, complete prompts, and item-by-item outputs were not made public.

  • System Card safety scores are not equivalent to ordinary task success rates; updated snapshots of earlier models also affect comparisons.

  • Direct and indirect injection results correspond only to the specified attack environment and cannot be used to infer the actual risk of every third-party tool.

Reproduction steps

  1. Define the instruction hierarchy for the system, developer, user, and tool outputs for each tool.

  2. Inject one privilege-escalating instruction separately through connectors, search, and function calls, and record whether the model changes the original task.

  3. Set confirmation thresholds separately for financial, high-risk, and general operations, and record correctness and refusal rates.

  4. Re-test the risk of overwriting changes using isolated files and adversarial modifications, and report the avoidance-only and correctness metrics.

  5. Write every privilege escalation, refusal, confirmation, and rollback event to an audit log.

Original evidence and data

  • The System Card states that GPT‑Red was added on 2026-08-03, indicating that the page is a continuously updated version.

  • OpenAI says Sol and Terra can find vulnerabilities and exploit snippets, but failed to complete autonomous end-to-end attacks against hardened targets.

  • The card also notes that higher-capability models are generally better than Terra/Luna at avoiding edit conflicts in complex tasks.

Scope boundaries

  • This is evidence about deployment safety, not a ranking of Sol’s code quality or writing quality.

  • “High” is a capability classification in the OpenAI Preparedness Framework, not a probability estimate of risk to ordinary users.

  • Every production deployment should be re-evaluated based on tool permissions, data sensitivity, and rollback capabilities, rather than copying isolated scores from the table.

Source excerpt or observation (compliance short quote only)

The official core explanation of the safety stack is “more than the sum of its parts,” emphasizing the joint role of the model, real-time monitoring, and account-level intervention.

What this supports

  • Supports discussing confirmation, injection, and authority boundaries for tool-using agents.
  • Supports retaining confirmation, audit, and rollback for high-risk writes.

What this does not support

  • Does not support converting safety scores into ordinary-task success or production-risk probabilities.
  • Tool permissions, data sensitivity, and deployment environment require retesting.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

OpenAI Deployment Safety Hub · OpenAI · Original publication date 2026-07-09 · Site edit date 2026-09-20

Open original source

GPT-5.6 Sol

Compare GPT-5.6 Sol in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Sol: Specs, Access, Changes, and the Risks That Still Matter

OpenAI's current GPT-5.6 Sol model page lists a 1.05M context window, 128K max output, reasoning controls, and a time-sensitive API price card. Here is what those facts mean for API, Codex, and browser users.

Related reviews

METR: Sol's time horizon changes with cheating treatmentIn Time Horizon 1.1 ReAct, METR estimates Sol's 50% time horizon at about 11.3 hours when cheating fails, over 270 hours when it succeeds, and about 71 hours when samples are dropped; none is robust.OpenAI release note: Sol's official results on long-horizon, coding, and knowledge workOpenAI reports Sol at 53.6 on Agents’ Last Exam, near Fable 5 on the Intelligence Index, and 80 on the Coding Agent Index, plus 92.2% on BrowseComp and 62.6% on OSWorld 2.0; these are dated vendor results.CodeRabbit: Sol's trade-offs in long coding-agent runs and code reviewCodeRabbit reports a 63.7% long-run coding pass rate for Sol with 20,968 average output tokens per completed task; review passed 69/99 actionable cases at 31.6% precision while producing 231 comments, combining recall gains with noise.Every: Sol excels as a collaborative knowledge-work partner, not as judgmentEvery describes Sol as fast and steerable across 24 drafts, email, meetings, and retrieval, but it scored 56/100 versus Fable's 90/100 on Senior Engineer and ranked last of six in writing; collaboration is not autonomous judgment.Deliver code with prediction, planning, review, and verificationSplit long-running coding into prediction, planning, implementation, adversarial review, and independent verification, checking the plan, tests, and stop conditions item by item; this is a commenter’s personal workflow, not Codex’s default configuration.Manage ChatGPT work with outcomes, context, and checkpointsReplace step-by-step micromanagement with an outcome, useful context, authorized actions, and review checkpoints; the author’s descriptions of Sol, Terra, Luna, and ChatGPT Work must be checked against the current product before use.Write Sol task prompts around outcomes and acceptance criteriaFrame research, coding, writing, and analysis requests with an outcome, success criteria, authority boundary, and validation method; the article’s insider attribution and internal comparison figures are not independently verified by this source.Design a verifiable multi-agent workflow with the Responses APISeparate judgment from deterministic processing, then combine programmatic tool calls, parallel subagents, and prompt-cache boundaries into a long-running workflow whose cost, latency, citations, and failures can be reviewed.