Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaGPT-6 Sol

GPT-6 Sol: Artificial Analysis on Cost Efficiency and Hallucination Measurement

Original source

Artificial Analysis

AuthorArtificial Analysis

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

Artificial Analysis's published results show that GPT-6 Sol max scores 48 on the Intelligence Index and 57 on the Coding Agent Index. The former costs about half as much per task as GPT-5.6 Sol max, while the Coding Agent Index score is 2 points higher. The AA-Omniscience hallucination rate falls from 92% to 60%, but accuracy also falls from 59% to 54%, and the share of questions Sol answers drops from 99% to 83%.

Test setup

  • Intelligence Index: v4.3, covering 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.

  • Coding Agent Index: v1.5, covering DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA. The article explicitly says this index uses the OpenAI Codex harness.

  • Reasoning level: The main comparisons use the max configurations for Sol and Luna.

  • AA-Omniscience methodology: The Index rewards correct answers and penalizes hallucinations, but does not penalize refusals. Scores range from -100 to 100; 0 means the number of correct answers equals the number of incorrect answers. Accuracy is the share of all questions answered completely correctly. Hallucination Rate is Incorrect / (Incorrect + Partial + Not Attempted).

  • GDPval-AA v2.1: The chart uses DeepSeek V4.1 Flash max as the 1600 Elo anchor. Scores are relative Elo ratings, not percentages.

Inputs and configuration

  • The article compares GPT-6 Sol max with GPT-5.6 Sol max, and also reports GPT-6 Luna max and GPT-5.6 Luna max for several metrics.

  • Prices per million tokens fall from $4/$20 to $2/$10 for Sol (input/output), and from $0.20/$1.20 to $0.10/$0.50 for Luna. The article says both generations have a 90% discount for cached reads and a 25% premium for cache writes.

  • Per-task Intelligence Index costs are based on weighted token consumption across AA evaluation tasks. The article says Sol averages about 31K output tokens per task, versus about 29K for the previous generation; Luna averages about 51K, versus about 41K for its predecessor.

  • The article does not disclose the full prompts, sampling parameters, number of repeated runs for each evaluation, or per-question run traces.

Results

Intelligence Index and cost

Model (max)Intelligence Index v4.3Cost per Index taskAverage output tokens per task
GPT-6 Sol48$1.06~31K
GPT-5.6 Sol47$1.99~29K
GPT-6 Luna37$0.07~51K
GPT-5.6 Luna37$0.18~41K

Coding Agent Index

Model (Codex harness, max)Index v1.5Cost per Coding Agent taskComponent results listed in the article
GPT-6 Sol57$2.99Terminal-Bench 4.0: 43%; SWE-Atlas-QnA: 58%
GPT-5.6 Sol55Not reported as an absolute valueTerminal-Bench 4.0: 37%; SWE-Atlas-QnA: 54%
GPT-6 Luna41About 40% of GPT-5.6 Luna's cost (absolute value not reported)SWE-Atlas-QnA: 44%; DeepSWE v1.1: 64%
GPT-5.6 Luna43Baseline for comparisonSWE-Atlas-QnA: 49%; DeepSWE v1.1: 66%

The article also reports Terminal-Bench 4.0 results within the Intelligence Index: 44% versus 40% for Sol, and 13% versus 12% for Luna. These are Intelligence Index results; the Coding Agent Index uses the Codex harness, so the values should be understood in the context of their respective evaluation setups and should not be combined into a single score.

AA-Omniscience and other component evaluations

MetricGPT-6 Sol maxGPT-5.6 Sol maxGPT-6 Luna maxGPT-5.6 Luna max
AA-Omniscience Index27221-10
Fully correct answer rate54%59%44%43%
Hallucination rate60%92%77%93%
Share of questions attempted83%99%The article says it answers fewer questions; no rate reportedNot reported
AutomationBench-AA62%60%53%50%
Terminal-Bench 4.0 (Intelligence Index)44%40%13%12%

The article also reports that GPT-6 Sol is about 100 Elo points behind GPT-5.6 Sol on GDPval-AA v2.1, while Luna is about 75 Elo points behind its predecessor. On AA-Briefcase v1.1, Luna is about 45 Elo points lower, while Sol is roughly even. The article does not give absolute scores for these evaluations in its main text.

Conclusions

  • Results reported directly by the article: Sol's Intelligence Index rises by 1 point, while its cost per task falls from $1.99 to $1.06. Its Coding Agent Index rises by 2 points to 57, with the article listing a cost of $2.99. Luna's Intelligence Index is unchanged as its cost falls from $0.18 to $0.07; its Coding Agent Index drops by 2 points to 41.

  • The article's interpretation of the hallucination results: Sol attempts fewer answers, and incorrect answers fall by about a quarter, reducing its hallucination rate; its accuracy also falls by 5 percentage points. Luna's accuracy is roughly unchanged, it answers fewer questions, and its hallucination rate falls from 93% to 77%. A lower hallucination rate alone does not show that knowledge accuracy has improved.

  • The article's interpretation of the knowledge-work regressions: The Artificial Analysis team says it manually inspected hundreds of model outputs and believes regressions on GDPval-AA and AA-Briefcase usually reflect lower presentation quality or deliverables missing rubric requirements. This is the team's attribution of the results, not a separately quantified causal metric in the table.

  • The cost-frontier claim: The authors interpret Sol's cost and Coding Agent Index combination as placing it on the Pareto frontier. This is an analytical conclusion based on the set of models they compared; it does not mean Sol is the lowest-cost choice for every model, price, or task distribution.

Limitations and review

  • The article publishes Artificial Analysis's aggregate measurements, but not the full test set, per-question prompts, run traces, statistical error, or number of repetitions for each evaluation. Readers cannot independently recompute the overall scores from this article.

  • Costs depend on token consumption in the organization's task set and prices at the time. Cache hits, retries, tool overhead, and task length in a real workload will change the bill. Halving input/output prices does not mean that every real task will cost exactly half as much.

  • Coding Agent Index results are tied to the Codex harness; they cannot be directly generalized to other agents, toolchains, or tool-free chat.

  • For review, compare the currently published AA Intelligence Index v4.3, Coding Agent Index v1.5, and AA-Omniscience results separately, preserving the date, reasoning level, and harness. To reproduce costs, use the same task set and record tokens, outcomes, and prices for each task. This article does not provide enough material to reproduce AA's original full experiment.

Review steps

  1. Confirm the versions cited in the article on Artificial Analysis: Intelligence Index v4.3, Coding Agent Index v1.5, and AA-Omniscience.

  2. Compare the GPT-6 and GPT-5.6 Sol/Luna configurations at the same max level, recording overall scores, component scores, answer rate, accuracy, hallucination rate, and task cost separately.

  3. For an independent cost retest, hold the harness, task set, and prices constant, and record input/output tokens, caching, retries, and completion status for each task. Such a retest compares only your own sample and cannot substitute for AA's unpublished original test set.

Scope of applicability

  • These data can be used to compare the listed AA versions, evaluations, and max configurations. They do not directly represent other reasoning levels, agent harnesses, or real business tasks.

  • Sol's lower hallucination rate comes with lower answer rate and accuracy. Work that requires broad coverage should also assess unanswered questions and error rates.

  • Artificial Analysis's explanation for the knowledge-work regressions comes from the team's manual inspection and interpretation; the article provides no independently verifiable causal experiment.

Primary evidence and data

The source page's main text directly reports prices, per-task costs, output tokens, Coding Agent Index scores, changes in component evaluations, and Omniscience metrics. The accompanying charts identify the 10 evaluations in Intelligence Index v4.3 and the 3 benchmarks in Coding Agent Index v1.5, and show accuracy, hallucination-rate, and Elo comparisons. This article records Terminal-Bench results from the two harnesses separately.

Source excerpt or observation (short quotation for context)

The original says that the two new versions show both “progress in some evaluations and regressions in others.” Therefore, price, a composite index, or hallucination rate alone is not enough to infer overall quality across all tasks.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Sol

Use and compare models in Tabbit

GPT-6 Sol

Related reviews

OfficialOpenAI2026-09-22

GPT-6 Sol: Official Benchmarks and Evaluation Boundaries

MediaAI IQ2026-09-22

GPT-6 Sol on AI IQ: Model Profile and Benchmark Coverage

CommunityKillSwitch-Bench

GPT-6 Sol: KillSwitch-Bench Adversarial Esoteric-Language Coding Agent Benchmark

MediaArtificial Analysis2026-09

GPT-6 Sol: Artificial Analysis Comparison Across Six Configurations

GPT-6 Sol

Related prompts

OfficialOpenAI Developers

GPT-6 Sol Official API Model Configuration

OfficialOpenAI Developers

OpenAI GPT-6 Family Prompting Guide

OfficialOpenAI official release notes2026-09-22

GPT-6 Prompt Caching Optimization Workflow

OfficialOpenAI Developers

OpenAI's Official GPT-6 Async Tool-Calling Workflow