Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
OfficialGPT-6 Luna

OpenAI's Official Release: GPT-6 Luna Benchmark Results and Cost Positioning

Original source

OpenAI / Introducing GPT-6 Sol and Luna

AuthorOpenAI

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

OpenAI positions GPT-6 Luna as a low-cost, high-throughput model and reports improvements over its predecessor on AutomationBench, DeepSWE, and OSWorld, but the public page does not provide complete per-task data or all the configurations needed for replication.

Test setup

  • Model/access point: GPT-6 Luna; the release page also reports results at different reasoning-effort levels. The DeepSWE result uses max effort.

  • Benchmarks: AutomationBench 1.0.6, DeepSWE v1.1, OSWorld 2.0 offline, and an internal factuality evaluation.

  • Pricing: $0.10 per million input tokens and $0.50 per million output tokens; these are the API prices listed on the release page.

  • Comparisons: Primarily GPT-5.6 Luna, with GPT-6 Sol, GPT-6 Astra, and Claude models included for some tasks.

Inputs and configuration

The release page describes AutomationBench as cross-application business workflows spanning 47 tools and covering sales, marketing, operations, support, finance, and HR. DeepSWE consists of long-horizon software engineering tasks in real codebases. OSWorld 2.0 covers everyday and professional computer-use workflows.

Results

  • AutomationBench: At high effort, GPT-6 Luna scores 5.4 percentage points above GPT-5.6 Luna. OpenAI reports 58% lower cost per task. The page does not give their exact absolute scores in the body text.

  • DeepSWE v1.1: GPT-6 Luna max scores 66.6%. OpenAI says this is comparable to Claude Opus 5 and Fable 5 at medium effort; the cost per task is 93% and 96% lower, respectively.

  • OSWorld 2.0 offline: GPT-6 Luna max exceeds GPT-5.6 Sol medium. OpenAI says its cost is about one-tenth as much.

  • Factuality: OpenAI reports a substantial improvement for Luna. At higher effort, it can reach GPT-5.6 Sol's level at about one-hundredth the cost. This result uses internal, de-identified ChatGPT conversations in which users had previously flagged a factual error.

  • OpenAI notes that this factuality dataset is not representative of typical usage and that the results are not controlled for answer length.

Conclusion

The official results support considering Luna for cost-sensitive coding agents, cross-application workflows, and computer-use tasks. Teams should test it on their own tasks with their own tools, prompts, effort levels, and acceptance criteria. The release page's relative cost figures should not be treated as fixed production costs.

Limitations

  • These are vendor-reported results and do not replace independent replication.

  • The release page does not publish complete benchmark prompts, per-task results, run counts, confidence intervals, or harness configurations for every evaluation.

  • The benchmarks use different effort levels, tools, and model comparisons; their scores are not directly comparable across benchmarks.

  • The reported Claude Fable 5.1 cost on AutomationBench excludes Opus 5 fallbacks used on about 40% of tasks. OpenAI explicitly notes that this understates the cost.

  • The OSWorld and internal factuality results come from specific dataset versions and selection methods. They should not be generalized to all desktop tasks or everyday questions.

Replication steps

  1. Fix the model snapshot, API provider, reasoning effort, tool permissions, and benchmark versions.

  2. Run AutomationBench, DeepSWE, and OSWorld repeatedly on the same task sets and harnesses, comparing against GPT-5.6 Luna.

  3. Save the original inputs, tool traces, per-task scores, completion status, token counts, latency, and actual costs.

  4. Report task success rate, cost, and human-rated deliverable quality separately. Do not combine percentages from different benchmarks into one ranking.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Luna

Use and compare models in Tabbit

GPT-6 Luna

Related reviews

MediaArtificial Analysis / GPT-6 Sol and Luna push the cost efficiency frontier2026-09-22

Artificial Analysis Independent Evaluation: GPT-6 Luna's Cost, Intelligence, and Coding Results

MediaArtificial Analysis / GPT-6 Luna: Release Intelligence, Performance & Price2026-09

Artificial Analysis Release Dashboard: GPT-6 Luna Performance, Cost, and Latency Across Six Effort Levels

CommunityReddit / r/codex2026-09-23

Reddit r/codex User Reports: GPT-6 Luna Coding Experience and Early Risks

CommunityReddit / r/codex2026-09-23

Reddit r/codex Discussion of Artificial Analysis Rankings: Luna's Rank and Subjective Impressions

GPT-6 Luna

Related prompts

OfficialOpenAI Developers

GPT-6 Model Family Prompting Starter Guide

OfficialOpenAI Developers

GPT-6 Luna API Model Configuration

OfficialOpenAI Developers

General Prompt Engineering Guide for the OpenAI API

OfficialOpenAI official announcement2026-09-22

GPT-6 Prompt Caching and Long-Running Agent Optimization Workflow