Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaGPT-6 Astra

GPT-6 Astra: Artificial Analysis Comparison of Intelligence, Speed, and Cost Across Six Reasoning Tiers

Original source

Artificial Analysis

AuthorArtificial Analysis

Source date2026-09

Tabbit curation2026-09-08

Read original

One-sentence conclusion

Artificial Analysis's same-page comparison across six tiers shows that Astra Max has the highest intelligence index but also the highest per-task cost, while Low has the lowest time to first token and cost, supporting routing by task value rather than uniform use of the highest tier.

Applicable scenarios

  • Suitable tasks: Preliminary routing among Astra's Low, Medium, High, XHigh, and Max tiers and the site's labeled Non-reasoning variant; comparing aggregate intelligence, generation speed, and per-task cost.

  • Unsuitable tasks: Using the aggregate index as a substitute for acceptance testing on a specific codebase, browser, legal, or medical task; the page data is not equivalent to an SLA.

  • Applicable model versions: The six GPT-6 Astra test configurations displayed on the Artificial Analysis page on 2026-09-08.

  • Applicable clients, Agents, or APIs: The page says that speed is measured from the first-party API; the specific request parameters, repetition count, and model snapshot require further verification against its methodology page.

  • Recommended reasoning tier and parameters: Start by testing Low for low-risk, latency-sensitive, or cost-sensitive tasks; test High/Max when a higher aggregate score is needed, and use task-level success rates to determine whether the additional cost is worthwhile.

Test environment, inputs, and configuration

  • Intelligence Index v4.3 is a composite of 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.

  • Per-task cost is a weighted average: for each evaluation, the prices of input, cache-hit, cache-write, reasoning, and answer tokens are divided by the number of tasks, then aggregated using the Intelligence Index weights.

  • Output speed is defined as generated tokens/s after the first API chunk is received; first-party models use first-party API performance.

  • The page does not provide per-question inputs, random seeds, repetition counts, confidence intervals, or failure logs in the visible body.

Raw data

Astra configurationIntelligence Index v4.3Output speedCost per Intelligence Index task
Max5361 tokens/s$3.26
XHigh5357 tokens/s$2.31
High5157 tokens/s$1.72
Medium5055 tokens/s$1.54
Low4653 tokens/s$0.82
Non-reasoning (page label)45Not shown$1.71

The page also shows that Low has the lowest time to first answer token, at 2.55 seconds; the six tiers differ in per-task cost by up to approximately 4×. In the further information table, the context for all six tiers is 1M, and the listed aggregate token pricing fields are all $7.7; this field must not be conflated with the “per-evaluation task cost” in the table above.

As an adjacent release comparison from the same site and version, the page lists a highest Intelligence value of 47 for GPT-5.6 Sol, 42 for Terra, and 38 for Luna; these are the highest values for each release series, not a tier-by-tier comparison at the same effort level.

Conclusions and applicable boundaries

Max adds 2 points over High, but per-task cost rises from $1.72 to $3.26; XHigh and Max both score 53 in the collected snapshot, while XHigh is faster and cheaper. This supports using High or XHigh as candidates first, then validating Max on the hardest tasks, rather than defaulting to the highest tier.

The page labels one configuration Non-reasoning, but the OpenAI official model guide explicitly stated on the same collection date that GPT-6 Astra does not support none as a reasoning effort. The two labels or test paths may differ; until Artificial Analysis discloses the exact API request behind that label, none should not be sent to the OpenAI API on this basis.

This is a dynamic leaderboard snapshot. During collection, the search cache showed scores different from those on the current page; the final figures use the data visible after directly opening the original page on 2026-09-08. Any subsequent review must record the collection date and Index version.

Reproduction steps

  1. Lock Intelligence Index v4.3, its 10 component evaluations, dataset versions, weights, and scoring scripts.

  2. Fix the OpenAI model snapshot, request endpoint, reasoning effort, context, cache strategy, maximum output, and retry rules.

  3. Run the same dataset for each tier and save per-question inputs, outputs, scores, input/cache-write/cache-read/reasoning/answer tokens, time to first token, and generation speed.

  4. Recalculate per-task cost using the weighted formula published on the page, and report the repetition count, mean, dispersion, and failure rate.

  5. Add success rates and human acceptance time on target business tasks, because similar aggregate indices do not imply the same production workflow cost.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

GPT-6 Astra

Use and compare models in Tabbit

GPT-6 Astra

Related reviews

OfficialOpenAI official release page

GPT-6 Astra: OpenAI's Official Five-Benchmark Comparison with GPT-5.6 Sol

OfficialReddit (r/codex) and comments on an OpenAI Codex GitHub issue2026-09-08

GPT-6 Astra: Reproducible Telemetry on Quota Consumption from 30-Second Subagent Polling

MediaCodeRabbit Blog2026-09-04

GPT-6 Astra: CodeRabbit's Cross-File Code Review and Cost Evaluation

CommunityReddit (r/codex)2026-09-08

GPT-6 Astra: Field Report on 100K+ LOC Long-Horizon Tasks, Reasoning Tiers, and Fast Mode

GPT-6 Astra

Related prompts

OfficialOpenAI Developers

OpenAI GPT-6 Astra Prompting and Configuration Guide

OfficialOpenAI Developers

OpenAI GPT-6 Astra API Model Configuration and Cost Boundaries

CommunityX2026-09-05

GPT-6 Astra: Guide to Cleaning Up Skills and AGENTS.md Instructions

CommunityReddit (r/codex)2026-09-05

GPT-6 Astra: Complete Codex Repository Instruction Audit Prompt