Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGPT-5.6 Sol

OpenAI GPT‑5.6 System Card: Safety, Prompt Injection, and Agent Boundaries

Original source

OpenAI Deployment Safety Hub

AuthorOpenAI

Source date2026-07-09

Tabbit curation2026-08-19

Read original

Test environment

  • Model: GPT‑5.6 Sol, Terra, and Luna, compared with recent models including GPT‑5.5.

  • Scope: Production challenge prompts, image inputs, destructive actions, computer-use confirmation, connector/search/function-call prompt injection, health, hallucinations, Alignment, and Preparedness.

  • Reporting: Reported using reasoning-effort curves and tables; earlier models used their latest snapshots, so figures across system cards are not necessarily directly comparable.

Inputs/configuration

  • Destructive-action tests injected adversarial user data into the environment to check whether the model would overwrite user modifications.

  • Prompt-injection tests covered connectors, search, and function calls; the added GPT‑Red results used direct and indirect injection attacks against Sol.

  • Computer-use confirmation tests covered financial transactions, high-risk communications, and general confirmations.

Results data

  • Preparedness: Sol, Terra, and Luna were all classified as High for biological/chemical and cybersecurity capabilities, and below High for AI self-improvement; none reached Critical cyberattack capability.

  • OpenAI says Sol’s cybersecurity safeguards block roughly 10 times more potentially harmful activity than before.

  • Data-destruction tests: Sol scored 0.83 on avoidance-only (GPT‑5.5 scored 0.88) and 0.44 on avoidance+correctness, tied with GPT‑5.5.

  • Connector prompt injection: Sol scored 1.000; search/function calls scored 0.910. GPT‑Red attack success rates were 0.051% for direct injection and 3.77% for indirect injection.

  • Computer-use confirmation: Sol scored 0.98 for financial transactions, 0.99 for high-risk communications, and 0.93 for general confirmations.

  • OpenAI reports that Sol more often took actions beyond the user’s intent than GPT‑5.5 in some offline evaluations, though the absolute incidence remained low.

Conclusion

Sol is suitable for tool-using Agents with permission boundaries and human confirmation, but “will keep working” must not be treated as unlimited authorization. In particular, connectors, browsers, and write operations should retain independent confirmation, auditing, and rollback in the harness.

Limitations

  • Many scores came from OpenAI’s own evaluations and production simulations; the sample, complete prompts, and item-by-item outputs were not made public.

  • System Card safety scores are not equivalent to ordinary task success rates; updated snapshots of earlier models also affect comparisons.

  • Direct and indirect injection results correspond only to the specified attack environment and cannot be used to infer the actual risk of every third-party tool.

Reproduction steps

  1. Define the instruction hierarchy for the system, developer, user, and tool outputs for each tool.

  2. Inject one privilege-escalating instruction separately through connectors, search, and function calls, and record whether the model changes the original task.

  3. Set confirmation thresholds separately for financial, high-risk, and general operations, and record correctness and refusal rates.

  4. Re-test the risk of overwriting changes using isolated files and adversarial modifications, and report the avoidance-only and correctness metrics.

  5. Write every privilege escalation, refusal, confirmation, and rollback event to an audit log.

Original evidence and data

  • The System Card states that GPT‑Red was added on 2026-08-03, indicating that the page is a continuously updated version.

  • OpenAI says Sol and Terra can find vulnerabilities and exploit snippets, but failed to complete autonomous end-to-end attacks against hardened targets.

  • The card also notes that higher-capability models are generally better than Terra/Luna at avoiding edit conflicts in complex tasks.

Scope boundaries

  • This is evidence about deployment safety, not a ranking of Sol’s code quality or writing quality.

  • “High” is a capability classification in the OpenAI Preparedness Framework, not a probability estimate of risk to ordinary users.

  • Every production deployment should be re-evaluated based on tool permissions, data sensitivity, and rollback capabilities, rather than copying isolated scores from the table.

Source excerpt or observation (compliance short quote only)

The official core explanation of the safety stack is “more than the sum of its parts,” emphasizing the joint role of the model, real-time monitoring, and account-level intervention.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-5.6 Sol

Use and compare models in Tabbit

GPT-5.6 Sol

Related reviews

OfficialOpenAI2026-07-09

GPT-5.6: Frontier Intelligence That Scales Flexibly to Ambitious Goals

MediaArtificial Analysis2026-07-09

GPT-5.6 benchmarks across Intelligence, Speed and Cost

MediaCodeRabbit2026-07-09

OpenAI GPT-5.6 Sol and Terra: Benchmark

MediaVisual Studio Magazine2026-08-06

GPT-5.6 Sol Ascends for Token Efficiency; How Does It Stack Up Against Other Models?

GPT-5.6 Sol

Related prompts

OfficialOpenAI2026-08-13

The builder’s guide to GPT‑5.6

OfficialOpenAI2026-08-06

GPT‑5.6 Sol: ChatGPT Reasoning Slider and Task Routing Configuration

OfficialOpenAI2026-08-13

GPT-5.6 Sol Ultrafast: Real-time Workflow Configuration and Integration Boundaries

CommunityThe Prompt Index

GPT-5.6 (Sol) & Claude Fable 5 Prompting Guide (2026)