Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

GPT-5.6 Terra · Media / benchmark · Independent measurement

GPT-5.6 Terra: Artificial Analysis Intelligence, Cost, and Coding Agent Indices

Artificial Analysis places GPT-5.6 Terra's Intelligence Index, Coding Agent Index, and cost position in one comparison frame for cost-capability screening.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Source-specific observation
The 2026-07-09 page reports Terra max Intelligence Index 55 and Coding Agent Index 77 and positions it against Sol and Luna on cost and capability.
Published conditions
The indices and prices use Artificial Analysis's published harness and dated snapshot; they are not every OpenAI API tier or a user's bill.

Key data and applicable tasks

One-sentence takeaway

Artificial Analysis's indices give Terra max an Intelligence Index score of 55 and a Coding Agent Index score of 77, placing it between Sol and Luna on cost; however, Terra is not the optimal point on every cost/intelligence Pareto frontier.

Test environment

  • Indices: Artificial Analysis Intelligence Index v4.1 and Coding Agent Index.

  • Coding agent harness: The article says the index combines agent harnesses including Codex, Claude Code, and Grok Build, and covers DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.

  • Reasoning level: The article primarily reports the max configuration for each GPT-5.6 model and compares the cost/intelligence frontier across different reasoning efforts.

  • Data source: Artificial Analysis's unified indices and task-cost statistics; the article notes that it supported evaluating OpenAI's Sol, Terra, and Luna during the pre-release phase.

Inputs/configuration

The article discloses the index names, model configurations, per-task costs, and some component evaluations, but does not provide per-question prompts, complete sampling parameters, or downloadable Terra execution traces. Full reproduction requires Artificial Analysis's evaluation service or a public harness at the same version.

Results

  • Intelligence Index: Terra max scores 55; Sol max scores 59; Luna max scores 51.

  • Intelligence Index cost per task: approximately $0.55 for Terra, $1.04 for Sol, and $0.21 for Luna; the article summarizes Terra's and Luna's cost differences relative to Sol as approximately 50% and 80%, respectively.

  • Coding Agent Index: Terra max scores 77, Sol max scores 80, and Luna max scores 75.

  • Coding Agent cost: The article says that Terra max and Luna max are approximately 60% and 80% cheaper than Sol, respectively.

  • AA-Briefcase: The article mainly publishes Sol's comparison results and warns that the GPT-5.6 family still needs to be assessed separately on knowledge work using rubric, Elo, and Presentation Elo; Coding Agent Index should not be treated as a measure of professional document quality.

  • Pareto observation: The article says Luna and Sol continue to lead Terra on the Intelligence and cost frontiers; Terra max's “mid-range price” does not automatically make it the globally optimal value choice.

Conclusions

  • Terra is a clear mid-range cost/intelligence option: it is cheaper than Sol and has a higher intelligence score than Luna, but the optimal budget point cannot be determined from the model configuration alone.

  • For Codex- and terminal-based agents, Terra max's score of 77 is enough to place it in frontier coding comparisons; the practical choice still depends on the cost of task failures, the tool harness, and output length.

  • For batch workloads, first compare “per-task cost × pass rate” on your own task distribution rather than comparing only the price per million tokens.

Limitations

  • The article explicitly says that Artificial Analysis supported OpenAI during the pre-release phase, creating potential selection and configuration bias; it remains an external index, but should not be regarded as a completely conflict-free blind test.

  • The Coding Agent Index is tied to different agent harnesses; model scores and costs cannot be compared directly outside the context of the harness.

  • Per-task cost is measured for a specific date, reasoning level, and input distribution; caching, retries, tool fees, and account pricing will change the actual bill.

  • The article does not publish Terra's input, output, and failure cases for each question, so the overall score cannot be independently recalculated.

Reproduction steps

  1. Record the Artificial Analysis article version, model aliases, max configurations, and collection date.

  2. In the same coding agent harness, fix the task set, timeouts, tool version, and concurrency, then run Terra, Sol, and Luna separately.

  3. For each task, save the pass status, number of model turns, input/output tokens, tool calls, elapsed time, and dollar cost.

  4. Separate the results into Intelligence, Coding Agent, and Briefcase/document quality; do not apply cross-index weighting.

  5. Calculate each model's pass rate, cost per task, and cost per success, then compare them with the directional conclusions in the article.

Source excerpt or observation (for a compliant short quotation only)

The article reports “GPT-5.6 Terra (max) and Luna (max) score 55 and 51 respectively”. These figures are index scores, not percentages.

What this supports

  • It supports like-for-like index and cost comparison

What this does not support

  • It supports like-for-like index and cost comparison, not a specific repository success rate, future pricing, or Tabbit account availability.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis · Original publication date 2026-07-09 · Site edit date 2026-09-20

Open original source

GPT-5.6 Terra

Compare GPT-5.6 Terra in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

GPT-5.6 Terra: What It Is, Access, and Where It Fits

A sourced GPT-5.6 Terra overview covering API limits, Sol and Luna differences, access surfaces, cost boundaries, and practical risks.

Related reviews

GPT-5.6 Terra System Card: Safety Guardrails and Agent BoundariesThe OpenAI System Card places Terra safety results in concrete tool, sandbox, and prompt-injection tests; it supports boundary assessment, not a production defense guarantee.Official OpenAI GPT-5.6 Terra Benchmarks, Pricing, and Task BoundariesOpenAI's GPT-5.6 release places Terra within the Sol/Luna family and separates benchmarks, pricing, and task examples; it does not publish Terra business success rates.GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing BenchmarkThis evidence note records GPT-5.6 Terra Reddit LLMDevs Role-Based Few-Shot Routing Benchmark under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost CurveThis evidence note records Artificial Analysis: Positioning GPT-5.6 Terra on the Intelligence–Cost Curve under its published model, platform, date, and sample conditions; it is not a universal ranking or production guarantee.GPT-5.6 Terra Long-Context Cost Thresholds and Routing WorkflowDataCamp's Terra routing case uses input length, tool-call frequency, and terminal needs to route long-context work and budget the full request cost.GPT-5.6 Terra API Model Parameters and Tool ConfigurationThe OpenAI model page gives Terra's model ID, reasoning levels, context and output limits, and tool capabilities for pre-integration checks.GPT-5.6 Terra Frontend Interaction Prototype Prompts and Validation WorkflowOpenAI's release page shows short prompts for runnable frontend prototypes and makes browser rendering checks part of the iteration loop.Generating Entrance Animations and Layout Variations in Framer Agent with GPT-5.6 TerraTill Janek's Framer case combines a few design choices, design-system constraints, and page-level animation variants for Terra-led visual exploration.