OpenAI's tables place Astra ahead of GPT-5.6 Sol across the selected computer-use, professional-work, coding, and science benchmarks, while the footnotes show that harness, tools, cost tier, and dataset version determine how the numbers should be read.
Suitable tasks: Initial model screening for computer use, professional software, multi-step terminal engineering, CAD generation, and graduate-level scientific reasoning.
Unsuitable tasks: Treating official benchmark scores as an independent acceptance test, a production success rate, or a cost forecast for arbitrary tasks.
Applicable model versions: GPT-6 Astra and GPT-5.6 Sol; some tables include additional comparison models.
Applicable client, agent, or API: The evaluation settings described on the OpenAI page; OSWorld uses latency simulation, Terminal-Bench uses an agent/terminal harness, and BenchCAD is explicitly with tools.
Recommended reasoning tier and parameters: The page does not publish one unified effort, seed, or complete API configuration. Reproduction must align each benchmark's footnotes and model settings.
Agents' Last Exam: Complex professional tasks in real software; the page reports 59.3% for Astra and 53.6% for GPT-5.6 Sol. At the highest-scoring settings, Astra uses about 65% fewer output tokens than Opus 5.
OSWorld 2.0: The table labels v2026.08.08, offline set, and partial score; in latency simulations Astra scores 72.6% at about 40 minutes per task, versus Sol's 65.7% at about 75 minutes.
BenchCAD: With tools, models reconstruct 3D objects from multi-view renders by generating CAD code; geometric-overlap scores are 95.9% for Astra and 83.3% for Sol. OpenAI says the shown Astra configuration has about 43% lower estimated API cost than Sol.
Terminal-Bench 4.0: Complex terminal tasks spanning software engineering, system configuration, and data analysis; the table reports 57.9% for Astra and 37.3% for Sol, with about 9% lower estimated API cost per task for Astra.
GPQA Diamond: Graduate-level biology, chemistry, and physics reasoning; Astra scores 96.0% at the high-scoring setting. At a lower-cost setting Astra scores 94.9%, above Sol's best 94.6%, at about 37% lower estimated API cost.
| Benchmark (table or matching narrative setting) | GPT-6 Astra | GPT-5.6 Sol | Key conditions |
|---|---|---|---|
| Agents' Last Exam | 59.3% | 53.6% | Highest-scoring setting; professional software tasks |
| OSWorld 2.0 | 72.6% | 65.7% | v2026.08.08, offline set, partial score; about 40 vs 75 minutes/task |
| BenchCAD (with tools) | 95.9% | 83.3% | Geometric-overlap score; Astra estimated about 43% cheaper |
| Terminal-Bench 4.0 | 57.9% | 37.3% | Complex terminal tasks; Astra estimated about 9% cheaper per task |
| GPQA Diamond | 96.0% | 94.6% | High-scoring setting; lower-cost setting is Astra 94.9% vs Sol 94.6% |
The opening overview reports a different benchmark, Terminal-Bench Science 0.1: Astra 64.6% vs Fable 52.6%, with a lower-cost setting of Astra 61.1% vs Sol's best 22.4%. Those figures must not replace the values for the separate Terminal-Bench 4.0 benchmark in the table, 57.9% vs 37.3%. Preserve benchmark names, settings, and table precision.
The official evidence supports Astra's strong positioning in these public comparisons, especially BenchCAD, Terminal-Bench 4.0, and Agents' Last Exam. OSWorld combines score with a latency simulation, while GPQA separates high-score and lower-cost settings; a higher score and a faster or cheaper run are different dimensions. The page does not provide complete inputs, seeds, per-question outputs, failure logs, or independent reruns, and some tasks are internal or harness-specific. The cybersecurity and alignment sections also state that production safeguards change results, so unsafeguarded safety-evaluation scores should not be treated as ordinary product experience.
Lock benchmark versions, offline/online sets, tool permissions, harness, model snapshot, effort, token limits, and cost accounting for each benchmark.
Reproduce the five table items separately, recording score, task duration, input/output/reasoning tokens, tool steps, and failure types.
Run high-scoring and lower-cost settings separately; do not compare one setting's score with another setting's price.
For OSWorld report pass/partial score and per-task duration; for BenchCAD report geometric overlap and tool calls.
Reproduce GPT-5.6 Sol with the table's version and harness. Do not mix the Sol figure from Terminal-Bench Science 0.1 in the narrative with Terminal-Bench 4.0.
The page explicitly labels the OSWorld number offline set, partial score and separates Terminal-Bench Science from Terminal-Bench 4.0. Those labels determine that the figures from different sections cannot be merged into one overall ranking.
GPT-6 Astra