The OpenAI System Card places Terra safety results in concrete tool, sandbox, and prompt-injection tests; it supports boundary assessment, not a production defense guarantee.
OpenAI Deployment Safety Hub · Read evidenceGPT-5.6 Terra · Reviews and evidence
Which GPT-5.6 Terra conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
SonarSource reran Sol and Terra on 4,444 Java tasks, which informs static code-quality and security findings in that sample but does not replace repository validation.
SonarSource · Read evidenceOpenAI's GPT-5.6 release places Terra within the Sol/Luna family and separates benchmarks, pricing, and task examples; it does not publish Terra business success rates.
OpenAI official launch page · Read evidenceFull reviews and related reading
Selected evidence
GPT-5.6 Terra System Card: Safety Guardrails and Agent Boundaries
The OpenAI System Card places Terra safety results in concrete tool, sandbox, and prompt-injection tests; it supports boundary assessment, not a production defense guarantee.
- Source-specific observation
- The 2026-07-09 card, amended 2026-08-03, covers Production Benchmarks, data-destruction avoidance, computer-use confirmation, connector/search prompt injection, and cyber and bio evaluations.
- Published conditions
- The cyber section uses a headless Linux box, common offensive tools, and three rollout pass@1; ExploitBench uses five seeds, while full samples are not public.
GPT-5.6 Terra: SonarSource's Retest of Code Quality and Security on 4,444 Java Tasks
SonarSource reran Sol and Terra on 4,444 Java tasks, which informs static code-quality and security findings in that sample but does not replace repository validation.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The July 11, 2026 SonarSource rerun covers 4,444 Java tasks and compares Sol and Terra with SonarQube quality and security rules.
- Published conditions
- It is SonarSource's task set and analysis pipeline rather than a general end-to-end agent test; complete prompts, repeats, and every snapshot are not public.
Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task Boundaries
OpenAI's GPT-5.6 release places Terra within the Sol/Luna family and separates benchmarks, pricing, and task examples; it does not publish Terra business success rates.
- Source-specific observation
- The 2026-07-09 release, with a 2026-07-30 price update, compares GPT-5.6 task examples, reasoning settings, and API prices.
- Published conditions
- The examples come from release demonstrations and official evaluations; no uniform seed, complete input, or Terra-specific frontend success rate is published.
GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent Indices
Artificial Analysis places GPT-5.6 Terra's Intelligence Index, Coding Agent Index, and cost position in one comparison frame for cost-capability screening.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The 2026-07-09 page reports Terra max Intelligence Index 55 and Coding Agent Index 77 and positions it against Sol and Luna on cost and capability.
- Published conditions
- The indices and prices use Artificial Analysis's published harness and dated snapshot; they are not every OpenAI API tier or a user's bill.
All sources
All sources
GPT-5.6 Terra System Card: Safety Guardrails and Agent Boundaries
The OpenAI System Card places Terra safety results in concrete tool, sandbox, and prompt-injection tests; it supports boundary assessment, not a production defense guarantee.
- Source-specific observation
- The 2026-07-09 card, amended 2026-08-03, covers Production Benchmarks, data-destruction avoidance, computer-use confirmation, connector/search prompt injection, and cyber and bio evaluations.
- Published conditions
- The cyber section uses a headless Linux box, common offensive tools, and three rollout pass@1; ExploitBench uses five seeds, while full samples are not public.
GPT-5.6 Terra: SonarSource's Retest of Code Quality and Security on 4,444 Java Tasks
SonarSource reran Sol and Terra on 4,444 Java tasks, which informs static code-quality and security findings in that sample but does not replace repository validation.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The July 11, 2026 SonarSource rerun covers 4,444 Java tasks and compares Sol and Terra with SonarQube quality and security rules.
- Published conditions
- It is SonarSource's task set and analysis pipeline rather than a general end-to-end agent test; complete prompts, repeats, and every snapshot are not public.
Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task Boundaries
OpenAI's GPT-5.6 release places Terra within the Sol/Luna family and separates benchmarks, pricing, and task examples; it does not publish Terra business success rates.
- Source-specific observation
- The 2026-07-09 release, with a 2026-07-30 price update, compares GPT-5.6 task examples, reasoning settings, and API prices.
- Published conditions
- The examples come from release demonstrations and official evaluations; no uniform seed, complete input, or Terra-specific frontend success rate is published.
GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent Indices
Artificial Analysis places GPT-5.6 Terra's Intelligence Index, Coding Agent Index, and cost position in one comparison frame for cost-capability screening.
Unverified: the original source could not be rechecked.
- Source-specific observation
- The 2026-07-09 page reports Terra max Intelligence Index 55 and Coding Agent Index 77 and positions it against Sol and Luna on cost and capability.
- Published conditions
- The indices and prices use Artificial Analysis's published harness and dated snapshot; they are not every OpenAI API tier or a user's bill.
GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark
This evidence note records GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.
Unverified: the original source could not be rechecked.
- Model/version
- GPT-5.6 Terra; source date 2026-08-18; do not merge snapshots or reasoning tiers.
- Platform/harness
- Reddit r/LLMDevs; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark does not establish a universal rate beyond its published sample.
Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost Curve
This evidence note records Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost Curve under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.
Unverified: the original source could not be rechecked.
- Model/version
- GPT-5.6 Terra; source date 2026-07-11; do not merge snapshots or reasoning tiers.
- Platform/harness
- X / Official Artificial Analysis Account; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost Curve does not establish a universal rate beyond its published sample.
25+ Claude Code and Codex Unattended Agent Loops: Practical Batch Research with Terra
This evidence note records 25+ Claude Code and Codex Unattended Agent Loops: Practical Batch Research with Terra under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.
Unverified: the original source could not be rechecked.
- Model/version
- GPT-5.6 Terra; source date 2026-08-19; do not merge snapshots or reasoning tiers.
- Platform/harness
- Reddit / r/ClaudeCode; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; 25+ Claude Code and Codex Unattended Agent Loops: Practical Batch Research with Terra does not establish a universal rate beyond its published sample.
Four-Model Same-Prompt iPhone UI Blind Test: Evidence for Terra in Real-World Development Tasks
This evidence note records Four-Model Same-Prompt iPhone UI Blind Test: Evidence for Terra in Real-World Development Tasks under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.
Unverified: the original source could not be rechecked.
- Model/version
- GPT-5.6 Terra; source date 2026-07-10; do not merge snapshots or reasoning tiers.
- Platform/harness
- X; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; Four-Model Same-Prompt iPhone UI Blind Test: Evidence for Terra in Real-World Development Tasks does not establish a universal rate beyond its published sample.
Personal Quota Comparison Between Terra High and Luna Extra High in Codex
This evidence note records Personal Quota Comparison Between Terra High and Luna Extra High in Codex under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.
Unverified: the original source could not be rechecked.
- Model/version
- GPT-5.6 Terra; source date 2026-07-12; do not merge snapshots or reasoning tiers.
- Platform/harness
- X; the source-specific platform and harness remain the unit of observation.
- Sample/date boundary
- Collected 2026-08-20; Personal Quota Comparison Between Terra High and Luna Extra High in Codex does not establish a universal rate beyond its published sample.
GPT-5.6 Terra
Compare GPT-5.6 Terra in Tabbit
Model access, features, and permissions depend on your current client account.