Artificial Analysis's summary of six effort levels shows substantial differences in Luna's performance, speed, and cost per task: max scores highest, while low costs the least.
Models/access point: GPT-6 Luna low, medium, high, xhigh, max, and non-reasoning; the page lists OpenAI as the provider.
Aggregate benchmark: Artificial Analysis Intelligence Index v4.3.2. The page says it combines 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
Pricing basis: Cost per Intelligence Index task. Prices and task costs in the model leaderboard should be interpreted alongside the benchmark version and token usage.
Context window: The page lists 1M tokens.
The page provides aggregate scores, output speed, time to first token, cost per task, and the benchmark composition. It does not provide downloadable per-task inputs, all run configurations, or raw outputs.
| Effort | Intelligence Index | Output speed | Cost per Index task |
|---|---|---|---|
| max | 37 | 154 tokens/s | $0.07 |
| xhigh | 34 | 153 tokens/s | $0.04 |
| high | 32 | 144 tokens/s | $0.03 |
| medium | 29 | 143 tokens/s | $0.02 |
| low | 21 | 176 tokens/s | $0.0045 |
| non-reasoning | 18 | 136 tokens/s | $0.01 |
The page reports the lowest time to first token for non-reasoning, at 0.72 seconds.
The highest and lowest per-task costs differ by about 15x. Low is the fastest at output, but its aggregate index score is below those of higher effort levels.
Effort selection has a substantial effect on Luna's cost and aggregate performance. Low or non-reasoning may be worth evaluating for simple, high-volume tasks. For tasks requiring higher quality, test high or max and choose based on local acceptance metrics.
This is an interactive summary dashboard, not a complete benchmark methodology report.
The Index is a weighted aggregate across different evaluations; a single score can hide differences between tasks.
The page does not provide per-task data, variance, confidence intervals, or all API parameters for each effort level.
Speed, latency, and cost vary with provider, prompt length, output length, and service load.
Test all six effort levels with the same API provider, fixed inputs, and output-length requirements.
For each task, record completion status, human quality rating, input/output tokens, time to first token, total latency, and billed cost.
Compare results against the ten evaluation versions listed on the page; review task-level outcomes rather than treating the aggregate Index score as task acceptance.
Repeat each task multiple times and report the median and range of variation.
GPT-6 Luna