GPT-6 Luna review navigator
Official benchmarks, independent analysis, and community reports about GPT-6 Luna, clearly separated from Tabbit's own testing.
Official
1 source-checked resourcesMedia
2 source-checked resourcesArtificial Analysis Independent Evaluation: GPT-6 Luna's Cost, Intelligence, and Coding Results
Artificial Analysis's own evaluations show GPT-6 Luna max delivering results broadly similar to its predecessor at much lower task cost, while scoring slightly lower on the Coding Agent Index and showing knowledge-work deliverable quality issues on GDPval and .
Artificial Analysis Release Dashboard: GPT-6 Luna Performance, Cost, and Latency Across Six Effort Levels
Artificial Analysis's summary of six effort levels shows substantial differences in Luna's performance, speed, and cost per task: max scores highest, while low costs the least.
Community
2 source-checked resourcesReddit r/codex User Reports: GPT-6 Luna Coding Experience and Early Risks
Early Codex users reported both more concise answers and faster performance, as well as serious individual cases of code changes and violations of Agent instructions; the sample is too small and tasks were not standardized, so it cannot establish a model ranki.
Reddit r/codex Discussion of Artificial Analysis Rankings: Luna's Rank and Subjective Impressions
The discussion reflects disagreement over the credibility of aggregate leaderboards, price cuts, and real-world usage time. It can serve as a checklist of model-selection questions, but not as an independent evaluation.
GPT-6 Luna
Use and compare models in Tabbit
Official benchmarks, independent analysis, and community reports about GPT-6 Luna, clearly separated from Tabbit's own testing.