Model: GPT‑5.6 Sol, Terra, and Luna, compared with recent models including GPT‑5.5.
Scope: Production challenge prompts, image inputs, destructive actions, computer-use confirmation, connector/search/function-call prompt injection, health, hallucinations, Alignment, and Preparedness.
Reporting: Reported using reasoning-effort curves and tables; earlier models used their latest snapshots, so figures across system cards are not necessarily directly comparable.
Destructive-action tests injected adversarial user data into the environment to check whether the model would overwrite user modifications.
Prompt-injection tests covered connectors, search, and function calls; the added GPT‑Red results used direct and indirect injection attacks against Sol.
Computer-use confirmation tests covered financial transactions, high-risk communications, and general confirmations.
Preparedness: Sol, Terra, and Luna were all classified as High for biological/chemical and cybersecurity capabilities, and below High for AI self-improvement; none reached Critical cyberattack capability.
OpenAI says Sol’s cybersecurity safeguards block roughly 10 times more potentially harmful activity than before.
Data-destruction tests: Sol scored 0.83 on avoidance-only (GPT‑5.5 scored 0.88) and 0.44 on avoidance+correctness, tied with GPT‑5.5.
Connector prompt injection: Sol scored 1.000; search/function calls scored 0.910. GPT‑Red attack success rates were 0.051% for direct injection and 3.77% for indirect injection.
Computer-use confirmation: Sol scored 0.98 for financial transactions, 0.99 for high-risk communications, and 0.93 for general confirmations.
OpenAI reports that Sol more often took actions beyond the user’s intent than GPT‑5.5 in some offline evaluations, though the absolute incidence remained low.
Sol is suitable for tool-using Agents with permission boundaries and human confirmation, but “will keep working” must not be treated as unlimited authorization. In particular, connectors, browsers, and write operations should retain independent confirmation, auditing, and rollback in the harness.
Many scores came from OpenAI’s own evaluations and production simulations; the sample, complete prompts, and item-by-item outputs were not made public.
System Card safety scores are not equivalent to ordinary task success rates; updated snapshots of earlier models also affect comparisons.
Direct and indirect injection results correspond only to the specified attack environment and cannot be used to infer the actual risk of every third-party tool.
Define the instruction hierarchy for the system, developer, user, and tool outputs for each tool.
Inject one privilege-escalating instruction separately through connectors, search, and function calls, and record whether the model changes the original task.
Set confirmation thresholds separately for financial, high-risk, and general operations, and record correctness and refusal rates.
Re-test the risk of overwriting changes using isolated files and adversarial modifications, and report the avoidance-only and correctness metrics.
Write every privilege escalation, refusal, confirmation, and rollback event to an audit log.
The System Card states that GPT‑Red was added on 2026-08-03, indicating that the page is a continuously updated version.
OpenAI says Sol and Terra can find vulnerabilities and exploit snippets, but failed to complete autonomous end-to-end attacks against hardened targets.
The card also notes that higher-capability models are generally better than Terra/Luna at avoiding edit conflicts in complex tasks.
This is evidence about deployment safety, not a ranking of Sol’s code quality or writing quality.
“High” is a capability classification in the OpenAI Preparedness Framework, not a probability estimate of risk to ordinary users.
Every production deployment should be re-evaluated based on tool permissions, data sensitivity, and rollback capabilities, rather than copying isolated scores from the table.
The official core explanation of the safety stack is “more than the sum of its parts,” emphasizing the joint role of the model, real-time monitoring, and account-level intervention.
GPT-5.6 Sol