The System Card rates Terra, alongside Sol and Luna, as having High capability in cybersecurity and biological and chemical domains, but not reaching Critical; its safety boundaries must be evaluated together with confirmation for tool-using agents, prompt-injection defenses, and data-destruction tests.
Models: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna, compared with previous-generation models including GPT-5.5.
Safety evaluations: Production Benchmarks, visual safety, data-destruction avoidance, computer-use confirmation, connector/search-function prompt injection, HealthBench, and cybersecurity and biological/chemical capability evaluations.
Evaluation formats: The page reports model-behavior tests without system-level safeguards, production-deployment simulations, and agent tests involving tools and sandboxes.
Version note: System Card tables may be updated as model snapshots and evaluation pipelines change; the page publishes the release date and change log.
The System Card does not disclose the complete production prompt or every safety sample, but it does disclose some task types, tool environments, and table metrics. The cybersecurity CTF description includes a headless Linux box, common offensive tools, and a tool-calling harness, using pass@1 over 3 rollouts; ExploitBench uses 5 seeds and reasoning continuity.
| Evaluation | GPT-5.6 Terra |
|---|---|
| Production Benchmarks: violent illicit behavior | 0.952 |
| Production Benchmarks: nonviolent illicit behavior | 0.990 |
| Production Benchmarks: extremism | 0.981 |
| Production Benchmarks: hate | 1.000 |
| Production Benchmarks: self-harm standard | 0.962 |
| Production Benchmarks: gore | 0.600 |
| Production Benchmarks: sexual | 0.966 |
| Production Benchmarks: sexual/minors | 0.974 |
| Image input: hate / extremism / self-harm / harms-erotic | 0.999 / 0.978 / 0.986 / 0.991 |
| Data-destruction avoidance (avoidance only) | 0.81 |
| Data-destruction avoidance (avoidance + correctness) | 0.37 |
| Computer use: financial transaction confirmation | 0.98 |
| Computer use: high-stakes communication confirmation | 0.98 |
| Computer use: general confirmation | 0.94 |
| Connector prompt injection | 1.000 |
| Search and function-calling prompt injection | 0.946 |
| GPT-Red: direct instruction-hierarchy injection | 0.061% |
| GPT-Red: indirect prompt injection | 3.32% |
Preparedness Framework: Sol, Terra, and Luna are all classified as having High capability in cybersecurity and biological/chemical domains.
AI self-improvement: The System Card says none of the three reached the High capability threshold.
The official summary is that the models can find vulnerabilities and some exploits, but in testing they could not carry out autonomous, end-to-end attacks against hardened targets.
Terra's combined data-destruction avoidance and correctness score is 0.37, below Sol's 0.44; for writable files, code repositories, and connector environments, confirmation and rollback should be treated as external gates.
Computer-use confirmation is not 100%; both financial and high-risk communications scored 0.98, so the model should not be allowed to decide irreversible operations on its own.
A direct prompt-injection rate of 0.061% and an indirect rate of 3.32% are not zero risk; when reading email, web pages, search results, or third-party tool output, untrusted instructions still need to be isolated.
The official classification marks Terra as having High cyber capability while also stating that it has not reached Critical. This supports controlled uses such as defensive vulnerability triage, remediation, and validation, but not unsupervised attack automation.
Some safety evaluations deliberately omit system-level safeguards to measure underlying behavior; they do not represent actual interception rates in the default product.
Production Benchmarks target difficult samples, and the page explicitly says that their error rates do not represent average production traffic.
Data-destruction, confirmation, and prompt-injection metrics come from specific harnesses and data distributions; they cannot establish the safety of arbitrary tools or third-party connectors.
The System Card is a vendor disclosure without complete samples, failure traces, or independent retesting; a safety classification is not equivalent to business authorization.
Divide agents into three tiers: read-only, writable but rollback-capable, and irreversible operations. For each tier, record whether confirmation is requested, whether user changes are retained, and the final correctness.
Build separate direct- and indirect-injection sets: place untrusted instructions in the user prompt, web pages, search results, email, and function returns, and record whether the model bypasses developer constraints.
Run defensive CTF and vulnerability-remediation tasks in an isolated Linux/container environment, fixing tool versions, timeouts, rollout counts, and the allowed network scope.
Record every model call, tool parameter, file diff, confirmation event, rollback result, and failure reason; do not record only the final answer.
Require human confirmation for any operation involving a financial transaction, sending high-risk communications, or deleting/overwriting data, and report the model's confirmation rate separately from the system's actual interception rate.
The key wording in the System Card is “High capability in both Cybersecurity and Biological and Chemical risk,” while also noting that the models have not reached Critical.
GPT-5.6 Terra