OpenAI positions Terra as a balanced tier for everyday workloads: official tables show notable gains over GPT-5.5 across coding, computer use, academic benchmarks, tool use, and long-context evaluation; however, these numbers represent vendor-reported results and cannot substitute for independent verification under an identical evaluation harness.
Models / Access points: GPT-5.6 Terra; OpenAI API, ChatGPT Work, Codex. The page also compares Sol, Luna, GPT-5.5, and other models.
Reasoning / Configuration: The page differentiates between settings like max and ultra; the Terra tables do not disclose itemized prompts, temperature, tool permissions, or trial counts.
Pricing: The launch page lists Terra at $2.50/1M input tokens and $15/1M output tokens; a page update note indicates a 20% price reduction on 2026-07-30, and the official X update post further specifies the post-reduction API pricing as $2 input and $12 output per 1M tokens.
Evaluation scope: Agents’ Last Exam, GDPval-AA v2, Artificial Analysis Intelligence/Coding Agent Index, SWE-Bench Pro, DeepSWE, Terminal-Bench 2.1, OSWorld 2.0, BrowseComp, GPQA Diamond, MRCR v2, Toolathlon, and others.
The launch page does not disclose the complete input prompts, sample sizes, sampling methods, trial counts, temperatures, tool definitions, or error breakdowns for each evaluation; the tables below therefore represent a verifiable ledger of official results rather than a fully reproducible benchmark suite.
| Benchmark | GPT-5.6 Terra | GPT-5.5 | Notes |
|---|---|---|---|
| Agents’ Last Exam | 50.4% | 46.9% | Long-horizon professional workflows across 55 domains; vendor-reported results |
| GDPval-AA v2 | 1,593 Elo | 1,493.7 Elo | Elo rating, not accuracy |
| Artificial Analysis Intelligence Index v4.1 | 55 | 54.8 | Page cites external index |
| Artificial Analysis Coding Agent Index v1.1 | 77.4 | 76.4 | Page cites external index |
| SWE-Bench Pro | 63.4% | 59.4% | Real-world repository engineering benchmark |
| DeepSWE v1.1 | 69.6% | 67.0% | Long-horizon software engineering |
| Terminal-Bench 2.1 | 87.4% | 85.6% | Command-line workflows |
| Benchmark | GPT-5.6 Terra | GPT-5.5 |
|---|---|---|
| OSWorld 2.0 | 50.2% | 47.5% |
| BrowseComp | 87.5% | 84.4% |
| GPQA Diamond | 92.9% | 93.6% |
| FrontierMath Tier 1–3 (v2) | 84.9% | 85.3% |
| AutomationBench | 15.2% | 12.9% |
| Toolathlon | 53.1% | 55.6% |
| OpenAI MRCR v2 8-needle 256K–512K | 89.6% | 81.5% |
| OpenAI MRCR v2 8-needle 512K–1M | 72.5% | 74.0% |
On OSWorld 2.0, official claims state that GPT-5.6 Sol scored 62.6%, outperforming Opus 4.8 while consuming 85% fewer output tokens; Terra's table value is 50.2%, indicating that promotional claims for Sol cannot be directly extrapolated to Terra.
Terra scores 57.7% on SEC-Bench Pro, 52.9% on ExploitBench, and 23.2% on ExploitGym, demonstrating strong cybersecurity capabilities that nonetheless should not bypass access controls or be treated as security clearance/authorization.
Terra scores 89.6% on OpenAI MRCR v2 in the 256K–512K range, but drops to 72.5% in the 512K–1M range; long-context performance varies across intervals and cannot be broadly characterized as "reliable across the entire 1M context."
The launch page provides a copyable Work prompt example: Create an interactive spirograph to explain how it works. This serves as an official showcase example, not a standalone benchmark input for Terra.
Well-suited for: Everyday coding, command-line engineering, browsing/computer operations, long-context retrieval, routine tool orchestration, and cost-constrained agent execution.
Requires escalation / human verification: Peak-complexity architectural planning, critical security or financial decisions, tasks demanding extreme rigor on difficult academic challenges like GPQA / FrontierMath Tier 4, and complex design deliverables requiring full visual review.
Model selection: Official data supports positioning Terra as a balanced successor candidate above GPT-5.5; however, whether it outperforms the more affordable Luna or the more capable Sol depends on empirical trade-offs across task cost, latency, and the penalty for verification failures.
All results originate from the OpenAI launch page—some listed as external indices—and lack public evaluation harnesses and raw item-by-item outputs.
Different benchmarks use distinct tools, time budgets, prompt structures, and scoring methodologies; percentages cannot be directly compared across different benchmark rows.
Pricing, model aliases, and available access points are subject to change; this note reflects only what was visible on the page as of the collection date.
https://developers.openai.com/api/docs/models/gpt-5.6-terra returned ERR_CONNECTION_CLOSED in Tabbit; search snippets or speculative details were deliberately avoided in drafting this document.
User takeover has been requested to inspect the tab in Tabbit; once access is restored, model IDs, context limits, and API parameters should be backfilled and annotated separately from this page's results.
Pin Terra's API snapshot, reasoning tier, temperature, tool schemas, context window length, and output limits.
Re-test coding tasks like SWE-Bench / Terminal-Bench first, followed by OSWorld / BrowseComp, GPQA, and MRCR long-context evaluations.
For each task, log success rates, p50/p95 latency, input/output tokens, tool call counts, costs, and human verification findings.
Maintain an identical harness benchmark against GPT-5.5, Luna, and Sol; document per-item failure modes, and do not conflate official tables with independent reproduction data.
GPT-5.6 Terra