Artificial Analysis's same-page comparison across six tiers shows that Astra Max has the highest intelligence index but also the highest per-task cost, while Low has the lowest time to first token and cost, supporting routing by task value rather than uniform use of the highest tier.
Suitable tasks: Preliminary routing among Astra's Low, Medium, High, XHigh, and Max tiers and the site's labeled Non-reasoning variant; comparing aggregate intelligence, generation speed, and per-task cost.
Unsuitable tasks: Using the aggregate index as a substitute for acceptance testing on a specific codebase, browser, legal, or medical task; the page data is not equivalent to an SLA.
Applicable model versions: The six GPT-6 Astra test configurations displayed on the Artificial Analysis page on 2026-09-08.
Applicable clients, Agents, or APIs: The page says that speed is measured from the first-party API; the specific request parameters, repetition count, and model snapshot require further verification against its methodology page.
Recommended reasoning tier and parameters: Start by testing Low for low-risk, latency-sensitive, or cost-sensitive tasks; test High/Max when a higher aggregate score is needed, and use task-level success rates to determine whether the additional cost is worthwhile.
Intelligence Index v4.3 is a composite of 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
Per-task cost is a weighted average: for each evaluation, the prices of input, cache-hit, cache-write, reasoning, and answer tokens are divided by the number of tasks, then aggregated using the Intelligence Index weights.
Output speed is defined as generated tokens/s after the first API chunk is received; first-party models use first-party API performance.
The page does not provide per-question inputs, random seeds, repetition counts, confidence intervals, or failure logs in the visible body.
| Astra configuration | Intelligence Index v4.3 | Output speed | Cost per Intelligence Index task |
|---|---|---|---|
| Max | 53 | 61 tokens/s | $3.26 |
| XHigh | 53 | 57 tokens/s | $2.31 |
| High | 51 | 57 tokens/s | $1.72 |
| Medium | 50 | 55 tokens/s | $1.54 |
| Low | 46 | 53 tokens/s | $0.82 |
| Non-reasoning (page label) | 45 | Not shown | $1.71 |
The page also shows that Low has the lowest time to first answer token, at 2.55 seconds; the six tiers differ in per-task cost by up to approximately 4×. In the further information table, the context for all six tiers is 1M, and the listed aggregate token pricing fields are all $7.7; this field must not be conflated with the “per-evaluation task cost” in the table above.
As an adjacent release comparison from the same site and version, the page lists a highest Intelligence value of 47 for GPT-5.6 Sol, 42 for Terra, and 38 for Luna; these are the highest values for each release series, not a tier-by-tier comparison at the same effort level.
Max adds 2 points over High, but per-task cost rises from $1.72 to $3.26; XHigh and Max both score 53 in the collected snapshot, while XHigh is faster and cheaper. This supports using High or XHigh as candidates first, then validating Max on the hardest tasks, rather than defaulting to the highest tier.
The page labels one configuration Non-reasoning, but the OpenAI official model guide explicitly stated on the same collection date that GPT-6 Astra does not support none as a reasoning effort. The two labels or test paths may differ; until Artificial Analysis discloses the exact API request behind that label, none should not be sent to the OpenAI API on this basis.
This is a dynamic leaderboard snapshot. During collection, the search cache showed scores different from those on the current page; the final figures use the data visible after directly opening the original page on 2026-09-08. Any subsequent review must record the collection date and Index version.
Lock Intelligence Index v4.3, its 10 component evaluations, dataset versions, weights, and scoring scripts.
Fix the OpenAI model snapshot, request endpoint, reasoning effort, context, cache strategy, maximum output, and retry rules.
Run the same dataset for each tier and save per-question inputs, outputs, scores, input/cache-write/cache-read/reasoning/answer tokens, time to first token, and generation speed.
Recalculate per-task cost using the weighted formula published on the page, and report the repetition count, mean, dispersion, and failure rate.
Add success rates and human acceptance time on target business tasks, because similar aggregate indices do not imply the same production workflow cost.
GPT-6 Astra