Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaGPT-6 Luna

Artificial Analysis Independent Evaluation: GPT-6 Luna's Cost, Intelligence, and Coding Results

Original source

Artificial Analysis / GPT-6 Sol and Luna push the cost efficiency frontier

AuthorArtificial Analysis

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

Artificial Analysis's own evaluations show GPT-6 Luna max delivering results broadly similar to its predecessor at much lower task cost, while scoring slightly lower on the Coding Agent Index and showing knowledge-work deliverable quality issues on GDPval and Briefcase.

Test setup

  • Models/access points: GPT-6 Luna max and GPT-5.6 Luna max; the article also reports other effort levels and models for comparison.

  • Benchmarks: Artificial Analysis Intelligence Index v4.3, Coding Agent Index, AA-Omniscience, AutomationBench-AA, Terminal-Bench 4.0, GDPval-AA v2.1, AA-Briefcase v1.1, SWE-Atlas-QnA, and DeepSWE v1.1.

  • Harness: The Coding Agent Index explicitly uses the OpenAI Codex harness; other tasks use the setup specific to each benchmark.

  • Cost basis: Cost per task and weighted cost for Artificial Analysis Intelligence Index tasks. Dollar amounts depend on the model prices and token usage at the time of the article's collection.

Inputs and configuration

The article provides several benchmark versions and some cost, token-usage, score, and task-result data. However, the article itself does not publish all per-task inputs, raw artifacts, or complete run configurations for each benchmark.

Results

  • Intelligence Index: The article says GPT-6 Luna max scores roughly level with GPT-5.6 Luna max. Luna's cost per task falls from $0.18 to $0.07, about 60% lower. Reported output tokens per task for the two Luna generations are about 41k and 51k, respectively.

  • Coding Agent Index: GPT-6 Luna max scores 41, two points below GPT-5.6 Luna max. SWE-Atlas-QnA falls from 49% to 44%, and DeepSWE v1.1 from 66% to 64%.

  • AA-Omniscience: Luna max's hallucination rate falls from 93% to 77%, while accuracy rises from about 43% to 44%. The article notes that Luna answers fewer questions.

  • AutomationBench-AA: Luna rises from 50% to 53%. Terminal-Bench 4.0 rises from 12% to 13%.

  • GDPval-AA v2.1: The article reports a decline of about 75 Elo for Luna; AA-Briefcase v1.1 falls by about 45 Elo. After reviewing hundreds of artifacts, the authors attribute the regressions to weaker presentation quality and deliverables that omit rubric elements.

Conclusion

Luna max's main advantage is lower cost; it is not an across-the-board improvement for coding agents or knowledge work. For tasks where completeness, output format, and quality criteria matter, inspect the deliverables and run independent acceptance checks instead of choosing a model based only on its aggregate score or price per token.

Limitations

  • Artificial Analysis designed or operates these metrics and task sets; they are not universal industry standards.

  • The article summarizes benchmark versions, some harness details, and results, but does not provide every per-task input and raw artifact, which limits reproducibility.

  • Reported scores depend on the model effort level, harness, output length, and pricing assumptions.

  • The Intelligence Index combines mixed results and cannot directly predict performance on an individual team's tasks. The article also reports both gains and regressions across benchmarks.

Replication steps

  1. Record the benchmark versions listed in the article, the GPT-6 Luna max and GPT-5.6 Luna max configurations, and the Codex harness version.

  2. Fix the task set, tools, and scoring rules for SWE-Atlas-QnA, DeepSWE, AutomationBench, Terminal-Bench, and knowledge-work tasks.

  3. Save each task's input and output, completion status, token count, cost, and human rating. Pay particular attention to missing deliverable requirements and presentation quality.

  4. Report success rate and cost for each benchmark separately; do not generalize the article's aggregate metrics to different harnesses.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Luna

Use and compare models in Tabbit

GPT-6 Luna

Related reviews

OfficialOpenAI / Introducing GPT-6 Sol and Luna2026-09-22

OpenAI's Official Release: GPT-6 Luna Benchmark Results and Cost Positioning

MediaArtificial Analysis / GPT-6 Luna: Release Intelligence, Performance & Price2026-09

Artificial Analysis Release Dashboard: GPT-6 Luna Performance, Cost, and Latency Across Six Effort Levels

CommunityReddit / r/codex2026-09-23

Reddit r/codex User Reports: GPT-6 Luna Coding Experience and Early Risks

CommunityReddit / r/codex2026-09-23

Reddit r/codex Discussion of Artificial Analysis Rankings: Luna's Rank and Subjective Impressions

GPT-6 Luna

Related prompts

OfficialOpenAI Developers

GPT-6 Model Family Prompting Starter Guide

OfficialOpenAI Developers

GPT-6 Luna API Model Configuration

OfficialOpenAI Developers

General Prompt Engineering Guide for the OpenAI API

OfficialOpenAI official announcement2026-09-22

GPT-6 Prompt Caching and Long-Running Agent Optimization Workflow