Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialGPT-6 Astra

GPT-6 Astra: OpenAI's Official Five-Benchmark Comparison with GPT-5.6 Sol

Original source

OpenAI official release page

AuthorOpenAI

Tabbit curation2026-09-08

Read original

One-sentence takeaway

OpenAI's tables place Astra ahead of GPT-5.6 Sol across the selected computer-use, professional-work, coding, and science benchmarks, while the footnotes show that harness, tools, cost tier, and dataset version determine how the numbers should be read.

Use cases

  • Suitable tasks: Initial model screening for computer use, professional software, multi-step terminal engineering, CAD generation, and graduate-level scientific reasoning.

  • Unsuitable tasks: Treating official benchmark scores as an independent acceptance test, a production success rate, or a cost forecast for arbitrary tasks.

  • Applicable model versions: GPT-6 Astra and GPT-5.6 Sol; some tables include additional comparison models.

  • Applicable client, agent, or API: The evaluation settings described on the OpenAI page; OSWorld uses latency simulation, Terminal-Bench uses an agent/terminal harness, and BenchCAD is explicitly with tools.

  • Recommended reasoning tier and parameters: The page does not publish one unified effort, seed, or complete API configuration. Reproduction must align each benchmark's footnotes and model settings.

Test environment and inputs/configuration

  • Agents' Last Exam: Complex professional tasks in real software; the page reports 59.3% for Astra and 53.6% for GPT-5.6 Sol. At the highest-scoring settings, Astra uses about 65% fewer output tokens than Opus 5.

  • OSWorld 2.0: The table labels v2026.08.08, offline set, and partial score; in latency simulations Astra scores 72.6% at about 40 minutes per task, versus Sol's 65.7% at about 75 minutes.

  • BenchCAD: With tools, models reconstruct 3D objects from multi-view renders by generating CAD code; geometric-overlap scores are 95.9% for Astra and 83.3% for Sol. OpenAI says the shown Astra configuration has about 43% lower estimated API cost than Sol.

  • Terminal-Bench 4.0: Complex terminal tasks spanning software engineering, system configuration, and data analysis; the table reports 57.9% for Astra and 37.3% for Sol, with about 9% lower estimated API cost per task for Astra.

  • GPQA Diamond: Graduate-level biology, chemistry, and physics reasoning; Astra scores 96.0% at the high-scoring setting. At a lower-cost setting Astra scores 94.9%, above Sol's best 94.6%, at about 37% lower estimated API cost.

Results data

Benchmark (table or matching narrative setting)GPT-6 AstraGPT-5.6 SolKey conditions
Agents' Last Exam59.3%53.6%Highest-scoring setting; professional software tasks
OSWorld 2.072.6%65.7%v2026.08.08, offline set, partial score; about 40 vs 75 minutes/task
BenchCAD (with tools)95.9%83.3%Geometric-overlap score; Astra estimated about 43% cheaper
Terminal-Bench 4.057.9%37.3%Complex terminal tasks; Astra estimated about 9% cheaper per task
GPQA Diamond96.0%94.6%High-scoring setting; lower-cost setting is Astra 94.9% vs Sol 94.6%

The opening overview reports a different benchmark, Terminal-Bench Science 0.1: Astra 64.6% vs Fable 52.6%, with a lower-cost setting of Astra 61.1% vs Sol's best 22.4%. Those figures must not replace the values for the separate Terminal-Bench 4.0 benchmark in the table, 57.9% vs 37.3%. Preserve benchmark names, settings, and table precision.

Conclusion and applicability boundaries

The official evidence supports Astra's strong positioning in these public comparisons, especially BenchCAD, Terminal-Bench 4.0, and Agents' Last Exam. OSWorld combines score with a latency simulation, while GPQA separates high-score and lower-cost settings; a higher score and a faster or cheaper run are different dimensions. The page does not provide complete inputs, seeds, per-question outputs, failure logs, or independent reruns, and some tasks are internal or harness-specific. The cybersecurity and alignment sections also state that production safeguards change results, so unsafeguarded safety-evaluation scores should not be treated as ordinary product experience.

Reproduction steps

  1. Lock benchmark versions, offline/online sets, tool permissions, harness, model snapshot, effort, token limits, and cost accounting for each benchmark.

  2. Reproduce the five table items separately, recording score, task duration, input/output/reasoning tokens, tool steps, and failure types.

  3. Run high-scoring and lower-cost settings separately; do not compare one setting's score with another setting's price.

  4. For OSWorld report pass/partial score and per-task duration; for BenchCAD report geometric overlap and tool calls.

  5. Reproduce GPT-5.6 Sol with the table's version and harness. Do not mix the Sol figure from Terminal-Bench Science 0.1 in the narrative with Terminal-Bench 4.0.

Source observation (short excerpt)

The page explicitly labels the OSWorld number offline set, partial score and separates Terminal-Bench Science from Terminal-Bench 4.0. Those labels determine that the figures from different sections cannot be merged into one overall ranking.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Astra

Use and compare models in Tabbit

GPT-6 Astra

Related reviews

OfficialReddit (r/codex) and comments on an OpenAI Codex GitHub issue2026-09-08

GPT-6 Astra: Reproducible Telemetry on Quota Consumption from 30-Second Subagent Polling

MediaCodeRabbit Blog2026-09-04

GPT-6 Astra: CodeRabbit's Cross-File Code Review and Cost Evaluation

CommunityReddit (r/codex)2026-09-08

GPT-6 Astra: Field Report on 100K+ LOC Long-Horizon Tasks, Reasoning Tiers, and Fast Mode

MediaArtificial Analysis2026-09

GPT-6 Astra: Artificial Analysis Comparison of Intelligence, Speed, and Cost Across Six Reasoning Tiers

GPT-6 Astra

Related prompts

OfficialOpenAI Developers

OpenAI GPT-6 Astra Prompting and Configuration Guide

OfficialOpenAI Developers

OpenAI GPT-6 Astra API Model Configuration and Cost Boundaries

CommunityX2026-09-05

GPT-6 Astra: Guide to Cleaning Up Skills and AGENTS.md Instructions

CommunityReddit (r/codex)2026-09-05

GPT-6 Astra: Complete Codex Repository Instruction Audit Prompt