Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGPT-5.6 Terra

Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task Boundaries

Original source

OpenAI official launch page

AuthorOpenAI

Source date2026-07-10

Tabbit curation2026-08-20

Read original

One-sentence takeaway

OpenAI positions Terra as a balanced tier for everyday workloads: official tables show notable gains over GPT-5.5 across coding, computer use, academic benchmarks, tool use, and long-context evaluation; however, these numbers represent vendor-reported results and cannot substitute for independent verification under an identical evaluation harness.

Test environment

  • Models / Access points: GPT-5.6 Terra; OpenAI API, ChatGPT Work, Codex. The page also compares Sol, Luna, GPT-5.5, and other models.

  • Reasoning / Configuration: The page differentiates between settings like max and ultra; the Terra tables do not disclose itemized prompts, temperature, tool permissions, or trial counts.

  • Pricing: The launch page lists Terra at $2.50/1M input tokens and $15/1M output tokens; a page update note indicates a 20% price reduction on 2026-07-30, and the official X update post further specifies the post-reduction API pricing as $2 input and $12 output per 1M tokens.

  • Evaluation scope: Agents’ Last Exam, GDPval-AA v2, Artificial Analysis Intelligence/Coding Agent Index, SWE-Bench Pro, DeepSWE, Terminal-Bench 2.1, OSWorld 2.0, BrowseComp, GPQA Diamond, MRCR v2, Toolathlon, and others.

Input and configuration

The launch page does not disclose the complete input prompts, sample sizes, sampling methods, trial counts, temperatures, tool definitions, or error breakdowns for each evaluation; the tables below therefore represent a verifiable ledger of official results rather than a fully reproducible benchmark suite.

Key results

Professional and coding

BenchmarkGPT-5.6 TerraGPT-5.5Notes
Agents’ Last Exam50.4%46.9%Long-horizon professional workflows across 55 domains; vendor-reported results
GDPval-AA v21,593 Elo1,493.7 EloElo rating, not accuracy
Artificial Analysis Intelligence Index v4.15554.8Page cites external index
Artificial Analysis Coding Agent Index v1.177.476.4Page cites external index
SWE-Bench Pro63.4%59.4%Real-world repository engineering benchmark
DeepSWE v1.169.6%67.0%Long-horizon software engineering
Terminal-Bench 2.187.4%85.6%Command-line workflows

Computer use, academic, tools, and long context

BenchmarkGPT-5.6 TerraGPT-5.5
OSWorld 2.050.2%47.5%
BrowseComp87.5%84.4%
GPQA Diamond92.9%93.6%
FrontierMath Tier 1–3 (v2)84.9%85.3%
AutomationBench15.2%12.9%
Toolathlon53.1%55.6%
OpenAI MRCR v2 8-needle 256K–512K89.6%81.5%
OpenAI MRCR v2 8-needle 512K–1M72.5%74.0%

Other boundary signals

  • On OSWorld 2.0, official claims state that GPT-5.6 Sol scored 62.6%, outperforming Opus 4.8 while consuming 85% fewer output tokens; Terra's table value is 50.2%, indicating that promotional claims for Sol cannot be directly extrapolated to Terra.

  • Terra scores 57.7% on SEC-Bench Pro, 52.9% on ExploitBench, and 23.2% on ExploitGym, demonstrating strong cybersecurity capabilities that nonetheless should not bypass access controls or be treated as security clearance/authorization.

  • Terra scores 89.6% on OpenAI MRCR v2 in the 256K–512K range, but drops to 72.5% in the 512K–1M range; long-context performance varies across intervals and cannot be broadly characterized as "reliable across the entire 1M context."

  • The launch page provides a copyable Work prompt example: Create an interactive spirograph to explain how it works. This serves as an official showcase example, not a standalone benchmark input for Terra.

Conclusions

  • Well-suited for: Everyday coding, command-line engineering, browsing/computer operations, long-context retrieval, routine tool orchestration, and cost-constrained agent execution.

  • Requires escalation / human verification: Peak-complexity architectural planning, critical security or financial decisions, tasks demanding extreme rigor on difficult academic challenges like GPQA / FrontierMath Tier 4, and complex design deliverables requiring full visual review.

  • Model selection: Official data supports positioning Terra as a balanced successor candidate above GPT-5.5; however, whether it outperforms the more affordable Luna or the more capable Sol depends on empirical trade-offs across task cost, latency, and the penalty for verification failures.

Limitations

  • All results originate from the OpenAI launch page—some listed as external indices—and lack public evaluation harnesses and raw item-by-item outputs.

  • Different benchmarks use distinct tools, time budgets, prompt structures, and scoring methodologies; percentages cannot be directly compared across different benchmark rows.

  • Pricing, model aliases, and available access points are subject to change; this note reflects only what was visible on the page as of the collection date.

Restricted source notes

  • https://developers.openai.com/api/docs/models/gpt-5.6-terra returned ERR_CONNECTION_CLOSED in Tabbit; search snippets or speculative details were deliberately avoided in drafting this document.

  • User takeover has been requested to inspect the tab in Tabbit; once access is restored, model IDs, context limits, and API parameters should be backfilled and annotated separately from this page's results.

Reproduction steps

  1. Pin Terra's API snapshot, reasoning tier, temperature, tool schemas, context window length, and output limits.

  2. Re-test coding tasks like SWE-Bench / Terminal-Bench first, followed by OSWorld / BrowseComp, GPQA, and MRCR long-context evaluations.

  3. For each task, log success rates, p50/p95 latency, input/output tokens, tool call counts, costs, and human verification findings.

  4. Maintain an identical harness benchmark against GPT-5.5, Luna, and Sol; document per-item failure modes, and do not conflate official tables with independent reproduction data.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-5.6 Terra

Use and compare models in Tabbit

GPT-5.6 Terra

Related reviews

OfficialOpenAI Deployment Safety Hub2026-07-09

GPT-5.6 Terra System Card: Safety Guardrails and Agent Boundaries

OfficialSonarSource2026-08-06

GPT-5.6 Terra: SonarSource's Retest of Code Quality and Security on 4,444 Java Tasks

MediaArtificial Analysis2026-07-09

GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent Indices

CommunityReddit r/LLMDevs2026-08-18

GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark

GPT-5.6 Terra

Related prompts

OfficialOpenAI Developers

GPT-5.6 Terra API Model Parameters and Tool Configuration

OfficialOfficial OpenAI release2026-07-09

GPT-5.6 Terra Frontend Interaction Prototype Prompts and Validation Workflow

MediaDataCamp2026-08-04

GPT-5.6 Terra Long-Context Cost Thresholds and Routing Workflow

CommunityX2026-08-01

Generating Entrance Animations and Layout Variations in Framer Agent with GPT-5.6 Terra