Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Terra · Official source · Vendor report

Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task Boundaries

OpenAI's GPT-5.6 release places Terra within the Sol/Luna family and separates benchmarks, pricing, and task examples; it does not publish Terra business success rates.

Official sourceVendor reportEdited 2026-09-20

Test conditions

Source-specific observation
The 2026-07-09 release, with a 2026-07-30 price update, compares GPT-5.6 task examples, reasoning settings, and API prices.
Published conditions
The examples come from release demonstrations and official evaluations; no uniform seed, complete input, or Terra-specific frontend success rate is published.

Key data and applicable tasks

One-sentence takeaway

OpenAI positions Terra as a balanced tier for everyday workloads: official tables show notable gains over GPT-5.5 across coding, computer use, academic benchmarks, tool use, and long-context evaluation; however, these numbers represent vendor-reported results and cannot substitute for independent verification under an identical evaluation harness.

Test environment

  • Models / Access points: GPT-5.6 Terra; OpenAI API, ChatGPT Work, Codex. The page also compares Sol, Luna, GPT-5.5, and other models.

  • Reasoning / Configuration: The page differentiates between settings like max and ultra; the Terra tables do not disclose itemized prompts, temperature, tool permissions, or trial counts.

  • Pricing: The launch page lists Terra at $2.50/1M input tokens and $15/1M output tokens; a page update note indicates a 20% price reduction on 2026-07-30, and the official X update post further specifies the post-reduction API pricing as $2 input and $12 output per 1M tokens.

  • Evaluation scope: Agents’ Last Exam, GDPval-AA v2, Artificial Analysis Intelligence/Coding Agent Index, SWE-Bench Pro, DeepSWE, Terminal-Bench 2.1, OSWorld 2.0, BrowseComp, GPQA Diamond, MRCR v2, Toolathlon, and others.

Input and configuration

The launch page does not disclose the complete input prompts, sample sizes, sampling methods, trial counts, temperatures, tool definitions, or error breakdowns for each evaluation; the tables below therefore represent a verifiable ledger of official results rather than a fully reproducible benchmark suite.

Key results

Professional and coding

BenchmarkGPT-5.6 TerraGPT-5.5Notes
Agents’ Last Exam50.4%46.9%Long-horizon professional workflows across 55 domains; vendor-reported results
GDPval-AA v21,593 Elo1,493.7 EloElo rating, not accuracy
Artificial Analysis Intelligence Index v4.15554.8Page cites external index
Artificial Analysis Coding Agent Index v1.177.476.4Page cites external index
SWE-Bench Pro63.4%59.4%Real-world repository engineering benchmark
DeepSWE v1.169.6%67.0%Long-horizon software engineering
Terminal-Bench 2.187.4%85.6%Command-line workflows

Computer use, academic, tools, and long context

BenchmarkGPT-5.6 TerraGPT-5.5
OSWorld 2.050.2%47.5%
BrowseComp87.5%84.4%
GPQA Diamond92.9%93.6%
FrontierMath Tier 1–3 (v2)84.9%85.3%
AutomationBench15.2%12.9%
Toolathlon53.1%55.6%
OpenAI MRCR v2 8-needle 256K–512K89.6%81.5%
OpenAI MRCR v2 8-needle 512K–1M72.5%74.0%

Other boundary signals

  • On OSWorld 2.0, official claims state that GPT-5.6 Sol scored 62.6%, outperforming Opus 4.8 while consuming 85% fewer output tokens; Terra's table value is 50.2%, indicating that promotional claims for Sol cannot be directly extrapolated to Terra.

  • Terra scores 57.7% on SEC-Bench Pro, 52.9% on ExploitBench, and 23.2% on ExploitGym, demonstrating strong cybersecurity capabilities that nonetheless should not bypass access controls or be treated as security clearance/authorization.

  • Terra scores 89.6% on OpenAI MRCR v2 in the 256K–512K range, but drops to 72.5% in the 512K–1M range; long-context performance varies across intervals and cannot be broadly characterized as "reliable across the entire 1M context."

  • The launch page provides a copyable Work prompt example: Create an interactive spirograph to explain how it works. This serves as an official showcase example, not a standalone benchmark input for Terra.

Conclusions

  • Well-suited for: Everyday coding, command-line engineering, browsing/computer operations, long-context retrieval, routine tool orchestration, and cost-constrained agent execution.

  • Requires escalation / human verification: Peak-complexity architectural planning, critical security or financial decisions, tasks demanding extreme rigor on difficult academic challenges like GPQA / FrontierMath Tier 4, and complex design deliverables requiring full visual review.

  • Model selection: Official data supports positioning Terra as a balanced successor candidate above GPT-5.5; however, whether it outperforms the more affordable Luna or the more capable Sol depends on empirical trade-offs across task cost, latency, and the penalty for verification failures.

Limitations

  • All results originate from the OpenAI launch page—some listed as external indices—and lack public evaluation harnesses and raw item-by-item outputs.

  • Different benchmarks use distinct tools, time budgets, prompt structures, and scoring methodologies; percentages cannot be directly compared across different benchmark rows.

  • Pricing, model aliases, and available access points are subject to change; this note reflects only what was visible on the page as of the collection date.

Restricted source notes

  • https://developers.openai.com/api/docs/models/gpt-5.6-terra returned ERR_CONNECTION_CLOSED in Tabbit; search snippets or speculative details were deliberately avoided in drafting this document.

  • User takeover has been requested to inspect the tab in Tabbit; once access is restored, model IDs, context limits, and API parameters should be backfilled and annotated separately from this page's results.

Reproduction steps

  1. Pin Terra's API snapshot, reasoning tier, temperature, tool schemas, context window length, and output limits.

  2. Re-test coding tasks like SWE-Bench / Terminal-Bench first, followed by OSWorld / BrowseComp, GPQA, and MRCR long-context evaluations.

  3. For each task, log success rates, p50/p95 latency, input/output tokens, tool call counts, costs, and human verification findings.

  4. Maintain an identical harness benchmark against GPT-5.5, Luna, and Sol; document per-item failure modes, and do not conflate official tables with independent reproduction data.

What this supports

  • It supports family positioning, parameter, and pricing boundaries

What this does not support

  • It supports family positioning, parameter, and pricing boundaries, not a blind-test ranking or current production cost guarantee.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

OpenAI official launch page · OpenAI · Original publication date 2026-07-10 · Site edit date 2026-09-20

Open original source

GPT-5.6 Terra

Compare GPT-5.6 Terra in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Terra: What It Is, Access, and Where It Fits

A sourced GPT-5.6 Terra overview covering API limits, Sol and Luna differences, access surfaces, cost boundaries, and practical risks.

Related reviews

GPT-5.6 Terra System Card: Safety Guardrails and Agent BoundariesThe OpenAI System Card places Terra safety results in concrete tool, sandbox, and prompt-injection tests; it supports boundary assessment, not a production defense guarantee.GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent IndicesArtificial Analysis places GPT-5.6 Terra's Intelligence Index, Coding Agent Index, and cost position in one comparison frame for cost-capability screening.GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing BenchmarkThis evidence note records GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost CurveThis evidence note records Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost Curve under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.GPT-5.6 Terra API Model Parameters and Tool ConfigurationThe OpenAI model page gives Terra's model ID, reasoning levels, context and output limits, and tool capabilities for pre-integration checks.GPT-5.6 Terra Frontend Interaction Prototype Prompts and Validation WorkflowOpenAI's release page shows short prompts for runnable frontend prototypes and makes browser rendering checks part of the iteration loop.GPT-5.6 Terra Long-Context Cost Thresholds and Routing WorkflowDataCamp's Terra routing case uses input length, tool-call frequency, and terminal needs to route long-context work and budget the full request cost.Generating Entrance Animations and Layout Variations in Framer Agent with GPT-5.6 TerraTill Janek's Framer case combines a few design choices, design-system constraints, and page-level animation variants for Terra-led visual exploration.