OpenAI positions GPT-6 Luna as a low-cost, high-throughput model and reports improvements over its predecessor on AutomationBench, DeepSWE, and OSWorld, but the public page does not provide complete per-task data or all the configurations needed for replication.
Model/access point: GPT-6 Luna; the release page also reports results at different reasoning-effort levels. The DeepSWE result uses max effort.
Benchmarks: AutomationBench 1.0.6, DeepSWE v1.1, OSWorld 2.0 offline, and an internal factuality evaluation.
Pricing: $0.10 per million input tokens and $0.50 per million output tokens; these are the API prices listed on the release page.
Comparisons: Primarily GPT-5.6 Luna, with GPT-6 Sol, GPT-6 Astra, and Claude models included for some tasks.
The release page describes AutomationBench as cross-application business workflows spanning 47 tools and covering sales, marketing, operations, support, finance, and HR. DeepSWE consists of long-horizon software engineering tasks in real codebases. OSWorld 2.0 covers everyday and professional computer-use workflows.
AutomationBench: At high effort, GPT-6 Luna scores 5.4 percentage points above GPT-5.6 Luna. OpenAI reports 58% lower cost per task. The page does not give their exact absolute scores in the body text.
DeepSWE v1.1: GPT-6 Luna max scores 66.6%. OpenAI says this is comparable to Claude Opus 5 and Fable 5 at medium effort; the cost per task is 93% and 96% lower, respectively.
OSWorld 2.0 offline: GPT-6 Luna max exceeds GPT-5.6 Sol medium. OpenAI says its cost is about one-tenth as much.
Factuality: OpenAI reports a substantial improvement for Luna. At higher effort, it can reach GPT-5.6 Sol's level at about one-hundredth the cost. This result uses internal, de-identified ChatGPT conversations in which users had previously flagged a factual error.
OpenAI notes that this factuality dataset is not representative of typical usage and that the results are not controlled for answer length.
The official results support considering Luna for cost-sensitive coding agents, cross-application workflows, and computer-use tasks. Teams should test it on their own tasks with their own tools, prompts, effort levels, and acceptance criteria. The release page's relative cost figures should not be treated as fixed production costs.
These are vendor-reported results and do not replace independent replication.
The release page does not publish complete benchmark prompts, per-task results, run counts, confidence intervals, or harness configurations for every evaluation.
The benchmarks use different effort levels, tools, and model comparisons; their scores are not directly comparable across benchmarks.
The reported Claude Fable 5.1 cost on AutomationBench excludes Opus 5 fallbacks used on about 40% of tasks. OpenAI explicitly notes that this understates the cost.
The OSWorld and internal factuality results come from specific dataset versions and selection methods. They should not be generalized to all desktop tasks or everyday questions.
Fix the model snapshot, API provider, reasoning effort, tool permissions, and benchmark versions.
Run AutomationBench, DeepSWE, and OSWorld repeatedly on the same task sets and harnesses, comparing against GPT-5.6 Luna.
Save the original inputs, tool traces, per-task scores, completion status, token counts, latency, and actual costs.
Report task success rate, cost, and human-rated deliverable quality separately. Do not combine percentages from different benchmarks into one ranking.
GPT-6 Luna