The launch page shows GPT-6 Sol improving on GPT-5.6 Sol in professional work, coding, and computer-use tasks, while approaching or exceeding some competitor results at a lower per-task cost. All scores and costs are figures published by OpenAI and should not be treated as independently reproduced results.
Models/effort levels: GPT-6 Sol and GPT-6 Luna; low, medium, high, xhigh, and max effort levels are reported by task.
AutomationBench 1.0.6: Uses 47 tools to measure end-to-end business workflows across sales, marketing, operations, customer support, finance, and HR.
Agents’ Last Exam V1: Long-horizon, economically valuable professional tasks spanning 55 subindustries.
FrontierCode 1.1 Main: Evaluates code correctness and mergeability, including test quality, scope control, coding style, and adherence to codebase conventions.
DeepSWE 1.1: Original, long-horizon software engineering tasks in real codebases.
OSWorld 2.0: Uses the offline-set partial reward from the v2026.08.08 release to test everyday and professional computer-use workflows.
OpenAI's GPT evaluations come from its research environments or API; system prompts and available tools may differ in the production ChatGPT product. Competitor data comes from public reports. Where Claude Fable 5 scores are shown without a specified effort level, the Fable 5.1 score is used.
| Benchmark/metric | GPT-6 Sol result | Official comparison and notes |
|---|---|---|
| AutomationBench 1.0.6 | xhigh: 33.2%, $0.27/task | Astra low: 30.3%, at 3.9× Sol's cost; Claude Opus 5 max: 26.9%, at 11.1× the cost; Claude Fable 5.1 + Opus 5 fallback max: 31.4%, at more than 8.9× the cost |
| Agents’ Last Exam V1 | max: 56.4% | Higher than Claude Opus 5's top score on this evaluation; Sol costs 60% less per task. The launch page does not list Opus 5's score or absolute cost. |
| FrontierCode 1.1 Main | The launch page says it is a substantial improvement over GPT-5.6 Sol. | The launch page says it matches Claude Fable 5.1 xhigh at far lower cost; the page body gives no exact score or cost. |
| DeepSWE 1.1 | max: 68.8% | Claude Fable 5 xhigh: 69.9%; Sol costs about 80% less per task. |
| OSWorld 2.0 offline partial reward | xhigh: 60.5% | Claude Opus 5 medium: 60.3%; Sol costs about 80% less per task. |
| Internal factuality evaluation | About half as many errors as GPT-5.6 Sol | OpenAI says reliability is close to Astra at substantially lower cost; the error-inducing conversation set does not represent everyday use. |
OpenAI also says Luna improved by 5.4 percentage points over GPT-5.6 Luna on AutomationBench at the high effort level, with 58% lower per-task cost; scored 66.6% on DeepSWE at max, close to Opus 5 and Fable 5 at medium, with costs 93% and 96% lower, respectively; and exceeded GPT-5.6 Sol at medium on OSWorld at max, at about one-tenth of its cost. For factuality, OpenAI also says Luna at higher effort can reach GPT-5.6 Sol's level at about one-hundredth of its cost.
| Model | Input price change | Output price change |
|---|---|---|
| GPT-5.6 Sol → GPT-6 Sol | $4 → $2 | $20 → $10 |
| GPT-5.6 Luna → GPT-6 Luna | $0.20 → $0.10 | $1.20 → $0.50 |
These results support considering GPT-6 Sol as a candidate model for long-horizon coding, business workflows, and computer-use agents, with per-task cost evaluated against the task budget. The tasks, harnesses, effort levels, and scoring methods differ across benchmarks, so their results cannot be combined into a single ranking. The cost advantages described on the launch page also do not imply equivalent savings on the same real-world business tasks.
All results were published by OpenAI. The page does not provide the full task samples, prompts, random seeds, per-item trajectories, and complete benchmark harness configurations needed for an independent rerun.
The AutomationBench data for Fable 5.1 + Opus 5 fallback omits Opus 5 fallback costs incurred on about 40% of tasks, so the page explicitly notes that its cost is understated.
The internal factuality test uses de-identified ChatGPT conversations that users had previously flagged as factually incorrect. It is a deliberately constructed error-inducing set and does not represent typical use. Scores were not controlled for answer length; OpenAI says its length sweep showed little impact.
The coding deception evaluation deliberately selects tasks that elicit dishonest behavior, with effort fixed at maximum. The launch page says this evaluation does not measure failure rates in everyday use. Do not interpret results from this kind of challenge set directly as everyday incidence rates.
GPT model evaluation environments may differ from the ChatGPT product due to differences in system prompts and tools. Competitor scores come from public reports, and comparison conditions are not necessarily consistent.
Costs and scores for AutomationBench, Agents’ Last Exam, FrontierCode, DeepSWE, and OSWorld should be understood in the context of each benchmark's tasks and version. The OSWorld result here is specifically the partial reward on the v2026.08.08 offline set.
Match the benchmark version, model effort level, tool permissions, and harness. If the launch page does not publish a configuration, list the missing details as barriers to reproduction.
For representative business tasks, record inputs, tool trajectories, outputs, tokens, task cost, elapsed time, and amount of human revision; run multiple trials and report variation.
Report scores and costs separately; do not present the official relative cost ratios as measured results for a local deployment or your own workflow.
The launch page body gives explicit numbers for Agents’ Last Exam, DeepSWE, OSWorld, factuality, and cross-benchmark cost comparisons; its AutomationBench chart provides effort levels, scores, and task-cost data. The FrontierCode body gives only a relative performance description, with no score that can be transcribed directly. This article preserves those disclosure boundaries and does not fill in unpublished values.
GPT-6 Sol