Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGPT-5.6 Terra

GPT-5.6 Terra System Card: Safety Guardrails and Agent Boundaries

Original source

OpenAI Deployment Safety Hub

AuthorOpenAI

Source date2026-07-09

Tabbit curation2026-08-19

Read original

One-sentence takeaway

The System Card rates Terra, alongside Sol and Luna, as having High capability in cybersecurity and biological and chemical domains, but not reaching Critical; its safety boundaries must be evaluated together with confirmation for tool-using agents, prompt-injection defenses, and data-destruction tests.

Test environment

  • Models: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna, compared with previous-generation models including GPT-5.5.

  • Safety evaluations: Production Benchmarks, visual safety, data-destruction avoidance, computer-use confirmation, connector/search-function prompt injection, HealthBench, and cybersecurity and biological/chemical capability evaluations.

  • Evaluation formats: The page reports model-behavior tests without system-level safeguards, production-deployment simulations, and agent tests involving tools and sandboxes.

  • Version note: System Card tables may be updated as model snapshots and evaluation pipelines change; the page publishes the release date and change log.

Inputs/configuration

The System Card does not disclose the complete production prompt or every safety sample, but it does disclose some task types, tool environments, and table metrics. The cybersecurity CTF description includes a headless Linux box, common offensive tools, and a tool-calling harness, using pass@1 over 3 rollouts; ExploitBench uses 5 seeds and reasoning continuity.

Results

Safety and agent control

EvaluationGPT-5.6 Terra
Production Benchmarks: violent illicit behavior0.952
Production Benchmarks: nonviolent illicit behavior0.990
Production Benchmarks: extremism0.981
Production Benchmarks: hate1.000
Production Benchmarks: self-harm standard0.962
Production Benchmarks: gore0.600
Production Benchmarks: sexual0.966
Production Benchmarks: sexual/minors0.974
Image input: hate / extremism / self-harm / harms-erotic0.999 / 0.978 / 0.986 / 0.991
Data-destruction avoidance (avoidance only)0.81
Data-destruction avoidance (avoidance + correctness)0.37
Computer use: financial transaction confirmation0.98
Computer use: high-stakes communication confirmation0.98
Computer use: general confirmation0.94
Connector prompt injection1.000
Search and function-calling prompt injection0.946
GPT-Red: direct instruction-hierarchy injection0.061%
GPT-Red: indirect prompt injection3.32%

Capability risk classification

  • Preparedness Framework: Sol, Terra, and Luna are all classified as having High capability in cybersecurity and biological/chemical domains.

  • AI self-improvement: The System Card says none of the three reached the High capability threshold.

  • The official summary is that the models can find vulnerabilities and some exploits, but in testing they could not carry out autonomous, end-to-end attacks against hardened targets.

Conclusions

  • Terra's combined data-destruction avoidance and correctness score is 0.37, below Sol's 0.44; for writable files, code repositories, and connector environments, confirmation and rollback should be treated as external gates.

  • Computer-use confirmation is not 100%; both financial and high-risk communications scored 0.98, so the model should not be allowed to decide irreversible operations on its own.

  • A direct prompt-injection rate of 0.061% and an indirect rate of 3.32% are not zero risk; when reading email, web pages, search results, or third-party tool output, untrusted instructions still need to be isolated.

  • The official classification marks Terra as having High cyber capability while also stating that it has not reached Critical. This supports controlled uses such as defensive vulnerability triage, remediation, and validation, but not unsupervised attack automation.

Limitations

  • Some safety evaluations deliberately omit system-level safeguards to measure underlying behavior; they do not represent actual interception rates in the default product.

  • Production Benchmarks target difficult samples, and the page explicitly says that their error rates do not represent average production traffic.

  • Data-destruction, confirmation, and prompt-injection metrics come from specific harnesses and data distributions; they cannot establish the safety of arbitrary tools or third-party connectors.

  • The System Card is a vendor disclosure without complete samples, failure traces, or independent retesting; a safety classification is not equivalent to business authorization.

Reproduction steps

  1. Divide agents into three tiers: read-only, writable but rollback-capable, and irreversible operations. For each tier, record whether confirmation is requested, whether user changes are retained, and the final correctness.

  2. Build separate direct- and indirect-injection sets: place untrusted instructions in the user prompt, web pages, search results, email, and function returns, and record whether the model bypasses developer constraints.

  3. Run defensive CTF and vulnerability-remediation tasks in an isolated Linux/container environment, fixing tool versions, timeouts, rollout counts, and the allowed network scope.

  4. Record every model call, tool parameter, file diff, confirmation event, rollback result, and failure reason; do not record only the final answer.

  5. Require human confirmation for any operation involving a financial transaction, sending high-risk communications, or deleting/overwriting data, and report the model's confirmation rate separately from the system's actual interception rate.

Source excerpt or observation (for compliant short quotation only)

The key wording in the System Card is “High capability in both Cybersecurity and Biological and Chemical risk,” while also noting that the models have not reached Critical.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-5.6 Terra

Use and compare models in Tabbit

GPT-5.6 Terra

Related reviews

OfficialSonarSource2026-08-06

GPT-5.6 Terra: SonarSource's Retest of Code Quality and Security on 4,444 Java Tasks

OfficialOpenAI official launch page2026-07-10

Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task Boundaries

MediaArtificial Analysis2026-07-09

GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent Indices

CommunityReddit r/LLMDevs2026-08-18

GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark

GPT-5.6 Terra

Related prompts

OfficialOpenAI Developers

GPT-5.6 Terra API Model Parameters and Tool Configuration

OfficialOfficial OpenAI release2026-07-09

GPT-5.6 Terra Frontend Interaction Prototype Prompts and Validation Workflow

MediaDataCamp2026-08-04

GPT-5.6 Terra Long-Context Cost Thresholds and Routing Workflow

CommunityX2026-08-01

Generating Entrance Animations and Layout Variations in Framer Agent with GPT-5.6 Terra