Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Terra · Official source · Vendor report

GPT-5.6 Terra: SonarSource's Retest of Code Quality and Security on 4,444 Java Tasks

SonarSource reran Sol and Terra on 4,444 Java tasks, which informs static code-quality and security findings in that sample but does not replace repository validation.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Source-specific observation
The July 11, 2026 SonarSource rerun covers 4,444 Java tasks and compares Sol and Terra with SonarQube quality and security rules.
Published conditions
It is SonarSource's task set and analysis pipeline rather than a general end-to-end agent test; complete prompts, repeats, and every snapshot are not public.

Key data and applicable tasks

One-sentence takeaway

On the same set of 4,444 Java tasks and with the medium reasoning configuration, Terra generated shorter code and had fewer missing completions, but its cognitive complexity and bug/vulnerability and code smell densities were higher, making static analysis and testing essential.

Test environment

  • Models: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.5.

  • Language: Java.

  • Number of tasks: 4,444, sourced from HumanEval, MBPP, and ComplexCodeEval.

  • Reasoning setting: medium for all three models.

  • Analyzer: SonarQube algorithmic code analysis.

  • Metric units: complexity/code smell per kLOC; bugs/vulnerabilities and category breakdowns per mLOC.

Inputs/configuration

SonarSource says the three models used the same tasks, the same quality analysis, and the same reasoning tier, and that GPT-5.5 was rerun in the same framework. The article does not disclose the prompt for each question, snapshots, sampling parameters, complete generated outputs, or a task-level failure list.

Results

MetricGPT-5.5GPT-5.6 SolGPT-5.6 Terra
Total lines of code702,720750,198617,132
Number of functions92,20682,16470,378
Comment share2.0%1.5%0.9%
Cognitive complexity / kLOC151.27143.23161.53
Functional skill pass rate78.66%81.99%79.96%
Missing completions0.27%0.25%0.18%
Bug density / mLOC504724763
Vulnerability density / mLOC68197203
Code smell density / kLOC17.0517.6023.31

Additional data: Terra produced approximately 8.37M output tokens, including approximately 3.66M reasoning tokens; input was approximately 1.32M tokens. Terra's blocker-level smell rate was 62/mLOC, lower than Sol's 72 and GPT-5.5's 78, but collection/generics issues rose to 11,843/mLOC. Concurrency/threading bugs were Terra's largest bug category, at approximately 350/mLOC.

Conclusions

  • Terra's 79.96% pass rate was slightly higher than GPT-5.5's but lower than Sol's 81.99%; production reliability cannot be judged by model price alone.

  • Terra completed the same task with 617,132 lines of code, substantially less than GPT-5.5 and Sol; this helps reduce review volume, but the problem density per kLOC was higher.

  • Cognitive complexity of 161.53/kLOC, bug density of 763/mLOC, vulnerability density of 203/mLOC, and smell density of 23.31/kLOC show that “shorter” does not mean “easier to maintain.”

  • Concurrency/threading, cryptographic configuration, and resource handling are risk areas that should receive priority in automated checks; manual read-through should not be the only gate.

Limitations

  • All tasks were synthetic/standardized Java tasks and cannot represent real multilingual repositories.

  • SonarQube static analysis reflects detectable code issues; it is not equivalent to the true online defect rate and cannot measure the business correctness of a complete system.

  • The same medium setting enables a controlled comparison, but the article does not tell us Terra's quality-cost curve at low/high/max settings.

  • SonarSource published the article and controlled the analyzer and metric definitions; retain the original outputs and cross-validate with a second testing framework.

Reproduction steps

  1. Fix the Java version, HumanEval/MBPP/ComplexCodeEval task versions, model aliases, and the medium reasoning tier.

  2. Save the prompt, response, compilation/test logs, generated code line count, and tokens for each task; do not save only aggregate scores.

  3. Analyze all three output groups with the same SonarQube version and rule set, calculating densities in the article's kLOC/mLOC units.

  4. Report pass rate, missing completions, and bugs/vulnerabilities/smells separately, broken down by severity.

  5. Manually sample concurrency, cryptography, and resource-handling issues to confirm whether static-rule hits constitute real defects.

Source excerpt or observation (for a compliant short quotation only)

The authors' summary of Terra is “Terra is the concise one”, but the same data shows that its cognitive complexity and problem density were the highest.

What this supports

  • It supports discussion of rule hits, defect repair, and security checks in this Java sample

What this does not support

  • It supports discussion of rule hits, defect repair, and security checks in this Java sample, not non-Java repositories, tool permissions, or production defect rates.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

SonarSource · Killian Carlsen-Phelan, Prasenjit Sarkar · Original publication date 2026-08-06 · Site edit date 2026-09-20

Open original source

GPT-5.6 Terra

Compare GPT-5.6 Terra in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Terra: What It Is, Access, and Where It Fits

A sourced GPT-5.6 Terra overview covering API limits, Sol and Luna differences, access surfaces, cost boundaries, and practical risks.

Related reviews

GPT-5.6 Terra System Card: Safety Guardrails and Agent BoundariesThe OpenAI System Card places Terra safety results in concrete tool, sandbox, and prompt-injection tests; it supports boundary assessment, not a production defense guarantee.Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task BoundariesOpenAI's GPT-5.6 release places Terra within the Sol/Luna family and separates benchmarks, pricing, and task examples; it does not publish Terra business success rates.GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent IndicesArtificial Analysis places GPT-5.6 Terra's Intelligence Index, Coding Agent Index, and cost position in one comparison frame for cost-capability screening.GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing BenchmarkThis evidence note records GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.GPT-5.6 Multi-Agent Routing Rules for Codex Agent.mdThis Codex configuration case routes Sol, Terra, and Luna by task complexity while retaining human or automated checkpoints for split work.Default Terra Sub-Agent Configuration for Codex Multi-Agent V2LPK's case uses Terra high as a bounded Codex Multi-Agent V2 sub-agent default and limits concurrency and depth to control capacity.GPT-5.6 Terra API Model Parameters and Tool ConfigurationThe OpenAI model page gives Terra's model ID, reasoning levels, context and output limits, and tool capabilities for pre-integration checks.GPT-5.6 Terra Frontend Interaction Prototype Prompts and Validation WorkflowOpenAI's release page shows short prompts for runnable frontend prototypes and makes browser rendering checks part of the iteration loop.