Artificial Analysis's published results show that GPT-6 Sol max scores 48 on the Intelligence Index and 57 on the Coding Agent Index. The former costs about half as much per task as GPT-5.6 Sol max, while the Coding Agent Index score is 2 points higher. The AA-Omniscience hallucination rate falls from 92% to 60%, but accuracy also falls from 59% to 54%, and the share of questions Sol answers drops from 99% to 83%.
Intelligence Index: v4.3, covering 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
Coding Agent Index: v1.5, covering DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA. The article explicitly says this index uses the OpenAI Codex harness.
Reasoning level: The main comparisons use the max configurations for Sol and Luna.
AA-Omniscience methodology: The Index rewards correct answers and penalizes hallucinations, but does not penalize refusals. Scores range from -100 to 100; 0 means the number of correct answers equals the number of incorrect answers. Accuracy is the share of all questions answered completely correctly. Hallucination Rate is Incorrect / (Incorrect + Partial + Not Attempted).
GDPval-AA v2.1: The chart uses DeepSeek V4.1 Flash max as the 1600 Elo anchor. Scores are relative Elo ratings, not percentages.
The article compares GPT-6 Sol max with GPT-5.6 Sol max, and also reports GPT-6 Luna max and GPT-5.6 Luna max for several metrics.
Prices per million tokens fall from $4/$20 to $2/$10 for Sol (input/output), and from $0.20/$1.20 to $0.10/$0.50 for Luna. The article says both generations have a 90% discount for cached reads and a 25% premium for cache writes.
Per-task Intelligence Index costs are based on weighted token consumption across AA evaluation tasks. The article says Sol averages about 31K output tokens per task, versus about 29K for the previous generation; Luna averages about 51K, versus about 41K for its predecessor.
The article does not disclose the full prompts, sampling parameters, number of repeated runs for each evaluation, or per-question run traces.
| Model (max) | Intelligence Index v4.3 | Cost per Index task | Average output tokens per task |
|---|---|---|---|
| GPT-6 Sol | 48 | $1.06 | ~31K |
| GPT-5.6 Sol | 47 | $1.99 | ~29K |
| GPT-6 Luna | 37 | $0.07 | ~51K |
| GPT-5.6 Luna | 37 | $0.18 | ~41K |
| Model (Codex harness, max) | Index v1.5 | Cost per Coding Agent task | Component results listed in the article |
|---|---|---|---|
| GPT-6 Sol | 57 | $2.99 | Terminal-Bench 4.0: 43%; SWE-Atlas-QnA: 58% |
| GPT-5.6 Sol | 55 | Not reported as an absolute value | Terminal-Bench 4.0: 37%; SWE-Atlas-QnA: 54% |
| GPT-6 Luna | 41 | About 40% of GPT-5.6 Luna's cost (absolute value not reported) | SWE-Atlas-QnA: 44%; DeepSWE v1.1: 64% |
| GPT-5.6 Luna | 43 | Baseline for comparison | SWE-Atlas-QnA: 49%; DeepSWE v1.1: 66% |
The article also reports Terminal-Bench 4.0 results within the Intelligence Index: 44% versus 40% for Sol, and 13% versus 12% for Luna. These are Intelligence Index results; the Coding Agent Index uses the Codex harness, so the values should be understood in the context of their respective evaluation setups and should not be combined into a single score.
| Metric | GPT-6 Sol max | GPT-5.6 Sol max | GPT-6 Luna max | GPT-5.6 Luna max |
|---|---|---|---|---|
| AA-Omniscience Index | 27 | 22 | 1 | -10 |
| Fully correct answer rate | 54% | 59% | 44% | 43% |
| Hallucination rate | 60% | 92% | 77% | 93% |
| Share of questions attempted | 83% | 99% | The article says it answers fewer questions; no rate reported | Not reported |
| AutomationBench-AA | 62% | 60% | 53% | 50% |
| Terminal-Bench 4.0 (Intelligence Index) | 44% | 40% | 13% | 12% |
The article also reports that GPT-6 Sol is about 100 Elo points behind GPT-5.6 Sol on GDPval-AA v2.1, while Luna is about 75 Elo points behind its predecessor. On AA-Briefcase v1.1, Luna is about 45 Elo points lower, while Sol is roughly even. The article does not give absolute scores for these evaluations in its main text.
Results reported directly by the article: Sol's Intelligence Index rises by 1 point, while its cost per task falls from $1.99 to $1.06. Its Coding Agent Index rises by 2 points to 57, with the article listing a cost of $2.99. Luna's Intelligence Index is unchanged as its cost falls from $0.18 to $0.07; its Coding Agent Index drops by 2 points to 41.
The article's interpretation of the hallucination results: Sol attempts fewer answers, and incorrect answers fall by about a quarter, reducing its hallucination rate; its accuracy also falls by 5 percentage points. Luna's accuracy is roughly unchanged, it answers fewer questions, and its hallucination rate falls from 93% to 77%. A lower hallucination rate alone does not show that knowledge accuracy has improved.
The article's interpretation of the knowledge-work regressions: The Artificial Analysis team says it manually inspected hundreds of model outputs and believes regressions on GDPval-AA and AA-Briefcase usually reflect lower presentation quality or deliverables missing rubric requirements. This is the team's attribution of the results, not a separately quantified causal metric in the table.
The cost-frontier claim: The authors interpret Sol's cost and Coding Agent Index combination as placing it on the Pareto frontier. This is an analytical conclusion based on the set of models they compared; it does not mean Sol is the lowest-cost choice for every model, price, or task distribution.
The article publishes Artificial Analysis's aggregate measurements, but not the full test set, per-question prompts, run traces, statistical error, or number of repetitions for each evaluation. Readers cannot independently recompute the overall scores from this article.
Costs depend on token consumption in the organization's task set and prices at the time. Cache hits, retries, tool overhead, and task length in a real workload will change the bill. Halving input/output prices does not mean that every real task will cost exactly half as much.
Coding Agent Index results are tied to the Codex harness; they cannot be directly generalized to other agents, toolchains, or tool-free chat.
For review, compare the currently published AA Intelligence Index v4.3, Coding Agent Index v1.5, and AA-Omniscience results separately, preserving the date, reasoning level, and harness. To reproduce costs, use the same task set and record tokens, outcomes, and prices for each task. This article does not provide enough material to reproduce AA's original full experiment.
Confirm the versions cited in the article on Artificial Analysis: Intelligence Index v4.3, Coding Agent Index v1.5, and AA-Omniscience.
Compare the GPT-6 and GPT-5.6 Sol/Luna configurations at the same max level, recording overall scores, component scores, answer rate, accuracy, hallucination rate, and task cost separately.
For an independent cost retest, hold the harness, task set, and prices constant, and record input/output tokens, caching, retries, and completion status for each task. Such a retest compares only your own sample and cannot substitute for AA's unpublished original test set.
These data can be used to compare the listed AA versions, evaluations, and max configurations. They do not directly represent other reasoning levels, agent harnesses, or real business tasks.
Sol's lower hallucination rate comes with lower answer rate and accuracy. Work that requires broad coverage should also assess unanswered questions and error rates.
Artificial Analysis's explanation for the knowledge-work regressions comes from the team's manual inspection and interpretation; the article provides no independently verifiable causal experiment.
The source page's main text directly reports prices, per-task costs, output tokens, Coding Agent Index scores, changes in component evaluations, and Omniscience metrics. The accompanying charts identify the 10 evaluations in Intelligence Index v4.3 and the 3 benchmarks in Coding Agent Index v1.5, and show accuracy, hallucination-rate, and Elo comparisons. This article records Terminal-Bench results from the two harnesses separately.
The original says that the two new versions show both “progress in some evaluations and regressions in others.” Therefore, price, a composite index, or hallucination rate alone is not enough to infer overall quality across all tasks.
GPT-6 Sol