Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
OfficialGPT-6 Sol

GPT-6 Sol: Official Benchmarks and Evaluation Boundaries

Original source

OpenAI

AuthorOpenAI

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

The launch page shows GPT-6 Sol improving on GPT-5.6 Sol in professional work, coding, and computer-use tasks, while approaching or exceeding some competitor results at a lower per-task cost. All scores and costs are figures published by OpenAI and should not be treated as independently reproduced results.

Test environment and methodology

  • Models/effort levels: GPT-6 Sol and GPT-6 Luna; low, medium, high, xhigh, and max effort levels are reported by task.

  • AutomationBench 1.0.6: Uses 47 tools to measure end-to-end business workflows across sales, marketing, operations, customer support, finance, and HR.

  • Agents’ Last Exam V1: Long-horizon, economically valuable professional tasks spanning 55 subindustries.

  • FrontierCode 1.1 Main: Evaluates code correctness and mergeability, including test quality, scope control, coding style, and adherence to codebase conventions.

  • DeepSWE 1.1: Original, long-horizon software engineering tasks in real codebases.

  • OSWorld 2.0: Uses the offline-set partial reward from the v2026.08.08 release to test everyday and professional computer-use workflows.

  • OpenAI's GPT evaluations come from its research environments or API; system prompts and available tools may differ in the production ChatGPT product. Competitor data comes from public reports. Where Claude Fable 5 scores are shown without a specified effort level, the Fable 5.1 score is used.

Results

Benchmark/metricGPT-6 Sol resultOfficial comparison and notes
AutomationBench 1.0.6xhigh: 33.2%, $0.27/taskAstra low: 30.3%, at 3.9× Sol's cost; Claude Opus 5 max: 26.9%, at 11.1× the cost; Claude Fable 5.1 + Opus 5 fallback max: 31.4%, at more than 8.9× the cost
Agents’ Last Exam V1max: 56.4%Higher than Claude Opus 5's top score on this evaluation; Sol costs 60% less per task. The launch page does not list Opus 5's score or absolute cost.
FrontierCode 1.1 MainThe launch page says it is a substantial improvement over GPT-5.6 Sol.The launch page says it matches Claude Fable 5.1 xhigh at far lower cost; the page body gives no exact score or cost.
DeepSWE 1.1max: 68.8%Claude Fable 5 xhigh: 69.9%; Sol costs about 80% less per task.
OSWorld 2.0 offline partial rewardxhigh: 60.5%Claude Opus 5 medium: 60.3%; Sol costs about 80% less per task.
Internal factuality evaluationAbout half as many errors as GPT-5.6 SolOpenAI says reliability is close to Astra at substantially lower cost; the error-inducing conversation set does not represent everyday use.

OpenAI also says Luna improved by 5.4 percentage points over GPT-5.6 Luna on AutomationBench at the high effort level, with 58% lower per-task cost; scored 66.6% on DeepSWE at max, close to Opus 5 and Fable 5 at medium, with costs 93% and 96% lower, respectively; and exceeded GPT-5.6 Sol at medium on OSWorld at max, at about one-tenth of its cost. For factuality, OpenAI also says Luna at higher effort can reach GPT-5.6 Sol's level at about one-hundredth of its cost.

API pricing (per million tokens)

ModelInput price changeOutput price change
GPT-5.6 Sol → GPT-6 Sol$4 → $2$20 → $10
GPT-5.6 Luna → GPT-6 Luna$0.20 → $0.10$1.20 → $0.50

Interpreting the results

These results support considering GPT-6 Sol as a candidate model for long-horizon coding, business workflows, and computer-use agents, with per-task cost evaluated against the task budget. The tasks, harnesses, effort levels, and scoring methods differ across benchmarks, so their results cannot be combined into a single ranking. The cost advantages described on the launch page also do not imply equivalent savings on the same real-world business tasks.

Limitations

  • All results were published by OpenAI. The page does not provide the full task samples, prompts, random seeds, per-item trajectories, and complete benchmark harness configurations needed for an independent rerun.

  • The AutomationBench data for Fable 5.1 + Opus 5 fallback omits Opus 5 fallback costs incurred on about 40% of tasks, so the page explicitly notes that its cost is understated.

  • The internal factuality test uses de-identified ChatGPT conversations that users had previously flagged as factually incorrect. It is a deliberately constructed error-inducing set and does not represent typical use. Scores were not controlled for answer length; OpenAI says its length sweep showed little impact.

  • The coding deception evaluation deliberately selects tasks that elicit dishonest behavior, with effort fixed at maximum. The launch page says this evaluation does not measure failure rates in everyday use. Do not interpret results from this kind of challenge set directly as everyday incidence rates.

  • GPT model evaluation environments may differ from the ChatGPT product due to differences in system prompts and tools. Competitor scores come from public reports, and comparison conditions are not necessarily consistent.

  • Costs and scores for AutomationBench, Agents’ Last Exam, FrontierCode, DeepSWE, and OSWorld should be understood in the context of each benchmark's tasks and version. The OSWorld result here is specifically the partial reward on the v2026.08.08 offline set.

Reproduction notes

  1. Match the benchmark version, model effort level, tool permissions, and harness. If the launch page does not publish a configuration, list the missing details as barriers to reproduction.

  2. For representative business tasks, record inputs, tool trajectories, outputs, tokens, task cost, elapsed time, and amount of human revision; run multiple trials and report variation.

  3. Report scores and costs separately; do not present the official relative cost ratios as measured results for a local deployment or your own workflow.

Original evidence and data

The launch page body gives explicit numbers for Agents’ Last Exam, DeepSWE, OSWorld, factuality, and cross-benchmark cost comparisons; its AutomationBench chart provides effort levels, scores, and task-cost data. The FrontierCode body gives only a relative performance description, with no score that can be transcribed directly. This article preserves those disclosure boundaries and does not fill in unpublished values.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Sol

Use and compare models in Tabbit

GPT-6 Sol

Related reviews

MediaArtificial Analysis2026-09-22

GPT-6 Sol: Artificial Analysis on Cost Efficiency and Hallucination Measurement

MediaAI IQ2026-09-22

GPT-6 Sol on AI IQ: Model Profile and Benchmark Coverage

CommunityKillSwitch-Bench

GPT-6 Sol: KillSwitch-Bench Adversarial Esoteric-Language Coding Agent Benchmark

MediaArtificial Analysis2026-09

GPT-6 Sol: Artificial Analysis Comparison Across Six Configurations

GPT-6 Sol

Related prompts

OfficialOpenAI Developers

GPT-6 Sol Official API Model Configuration

OfficialOpenAI Developers

OpenAI GPT-6 Family Prompting Guide

OfficialOpenAI official release notes2026-09-22

GPT-6 Prompt Caching Optimization Workflow

OfficialOpenAI Developers

OpenAI's Official GPT-6 Async Tool-Calling Workflow