Artificial Analysis's own evaluations show GPT-6 Luna max delivering results broadly similar to its predecessor at much lower task cost, while scoring slightly lower on the Coding Agent Index and showing knowledge-work deliverable quality issues on GDPval and Briefcase.
Models/access points: GPT-6 Luna max and GPT-5.6 Luna max; the article also reports other effort levels and models for comparison.
Benchmarks: Artificial Analysis Intelligence Index v4.3, Coding Agent Index, AA-Omniscience, AutomationBench-AA, Terminal-Bench 4.0, GDPval-AA v2.1, AA-Briefcase v1.1, SWE-Atlas-QnA, and DeepSWE v1.1.
Harness: The Coding Agent Index explicitly uses the OpenAI Codex harness; other tasks use the setup specific to each benchmark.
Cost basis: Cost per task and weighted cost for Artificial Analysis Intelligence Index tasks. Dollar amounts depend on the model prices and token usage at the time of the article's collection.
The article provides several benchmark versions and some cost, token-usage, score, and task-result data. However, the article itself does not publish all per-task inputs, raw artifacts, or complete run configurations for each benchmark.
Intelligence Index: The article says GPT-6 Luna max scores roughly level with GPT-5.6 Luna max. Luna's cost per task falls from $0.18 to $0.07, about 60% lower. Reported output tokens per task for the two Luna generations are about 41k and 51k, respectively.
Coding Agent Index: GPT-6 Luna max scores 41, two points below GPT-5.6 Luna max. SWE-Atlas-QnA falls from 49% to 44%, and DeepSWE v1.1 from 66% to 64%.
AA-Omniscience: Luna max's hallucination rate falls from 93% to 77%, while accuracy rises from about 43% to 44%. The article notes that Luna answers fewer questions.
AutomationBench-AA: Luna rises from 50% to 53%. Terminal-Bench 4.0 rises from 12% to 13%.
GDPval-AA v2.1: The article reports a decline of about 75 Elo for Luna; AA-Briefcase v1.1 falls by about 45 Elo. After reviewing hundreds of artifacts, the authors attribute the regressions to weaker presentation quality and deliverables that omit rubric elements.
Luna max's main advantage is lower cost; it is not an across-the-board improvement for coding agents or knowledge work. For tasks where completeness, output format, and quality criteria matter, inspect the deliverables and run independent acceptance checks instead of choosing a model based only on its aggregate score or price per token.
Artificial Analysis designed or operates these metrics and task sets; they are not universal industry standards.
The article summarizes benchmark versions, some harness details, and results, but does not provide every per-task input and raw artifact, which limits reproducibility.
Reported scores depend on the model effort level, harness, output length, and pricing assumptions.
The Intelligence Index combines mixed results and cannot directly predict performance on an individual team's tasks. The article also reports both gains and regressions across benchmarks.
Record the benchmark versions listed in the article, the GPT-6 Luna max and GPT-5.6 Luna max configurations, and the Codex harness version.
Fix the task set, tools, and scoring rules for SWE-Atlas-QnA, DeepSWE, AutomationBench, Terminal-Bench, and knowledge-work tasks.
Save each task's input and output, completion status, token count, cost, and human rating. Pay particular attention to missing deliverable requirements and presentation quality.
Report success rate and cost for each benchmark separately; do not generalize the article's aggregate metrics to different harnesses.
GPT-6 Luna