Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Opus 5.5

Artificial Analysis Evaluation: Claude Opus 5.5 Tops the Intelligence Index, with Cost and Output Measurements

Original source

Artificial Analysis

AuthorArtificial Analysis

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

In Artificial Analysis's snapshot dated 2026-09-22, Claude Opus 5.5 at the max effort level topped the rankings with an Intelligence Index score of 58. It was strong on knowledge work and several agent evaluations, but generated a large volume of output, with an estimated cost of about $5.98 per individual index task.

Use cases

  • Tasks this evaluation can help assess: General reasoning and knowledge, agentic knowledge work, terminal coding, scientific coding tasks, and automated workflows. It can also help compare index scores, output volume, and estimated task costs across effort levels.

  • Tasks these results should not be generalized to: Tasks outside the ten index evaluations, specific private enterprise workflows, or other reasoning parameters or agent configurations. AA-Briefcase uses a private evaluation set, so its results cannot be independently rerun based on the public summary alone.

  • Applicable model version: Claude Opus 5.5. The Artificial Analysis model page lists the tested variant as Adaptive Reasoning, with Anthropic's default fallback enabled.

  • Test environment or client: Artificial Analysis's own evaluations. AA-Briefcase uses its open-source reference agent harness, Stirrup. The article does not provide the full execution configuration, per-task inputs, or submission records for the other evaluations.

  • Reasoning levels and parameters: low, medium, high, xhigh, and max. The article says Anthropic's default fallback was enabled at all five levels for the Intelligence Index. It does not specify other reasoning parameters.

Evaluation method

Artificial Analysis Intelligence Index v4.3.2 aggregates ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The index version and evaluation list come from the same organization's model page; the rankings and costs are dynamic snapshots as visible on the collection date.

The article covers index evaluations at five effort levels. It specifically notes that AA-Briefcase v1.1 is Artificial Analysis's private frontier knowledge-work evaluation, which uses the open-source reference agent harness Stirrup to assess whether a model can produce accurate, well-presented professional deliverables. Artificial Analysis's methodology defines the cost of an individual task as a weighted average of input, cached, and output tokens, aggregated according to index weights. Cost therefore reflects both pricing and token consumption.

The article does not disclose the full prompts, sample counts for each evaluation, model call logs, or complete per-evaluation scores for every effort level. Except for AA-Briefcase, this note does not infer which specific agent harnesses the evaluations used.

Key results

  • Overall index: Claude Opus 5.5 max scored 58, the highest score at the time described in the article. The model page also showed it ranked #1 / 212 when collected. This ranking reflects the snapshot for v4.3.2 on that date and is not a permanent ranking.

  • Led in six evaluations: Humanity's Last Exam at 61.4% (previous high: Claude Fable 5.1 at 59.1%), SciCode at 66.9% (previous high: Fable 5.1 at 63.1%), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience, and AutomationBench-AA. The article does not provide complete figures for the other three evaluations.

  • Terminal and coding: Terminal-Bench 4.0 score was 59.6%, tied for the lead with GPT-6 Astra xhigh and 11 percentage points above Opus 5. The model still trailed others on CritPt, AA-LCR, and GDP.pdf.

  • Agentic knowledge work: AA-Briefcase v1.1 scored 1,822 Elo, 143 points above Fable 5.1. It led on the analysis-quality and presentation subscores, but scored slightly lower than Fable 5.1 on rubric grading. GDPval-AA v2.1 scored 1,846 Elo, 111 above Fable 5.1 and 138 above Opus 5.

  • Effort-level and cost frontier: Four of the five levels—max, xhigh, high, and medium—were on the Pareto frontier for Intelligence versus cost per individual task. The article describes the comparison set as models scoring above 50 on the index.

  • Output volume and per-task cost: The max level generated about 119k output tokens per index task, compared with about 73k for Opus 5 max, 78k for Fable 5.1 max, and 27k for GPT-6 Astra max. The model page's collection snapshot showed a weighted average cost of $5.98 per index task.

  • Total evaluation cost and cumulative output: The model page showed a total cost of $8,708.20 for the full Intelligence Index evaluation and 260M output tokens generated cumulatively. The cumulative total and the article's roughly 119k tokens per task use different measures and are not interchangeable.

  • Pricing and context: Artificial Analysis's page listed API prices of $4 per million input tokens, $20 per million output tokens, and $0.20 per million cache-read tokens; a 1M-token context window; text and image input; and text output. The page says prices are 20% lower than Opus 5's $5/$25 input/output prices, with cache reads down from $0.50 to $0.20. These are API prices listed on the page, not model capabilities measured by an independent benchmark.

Raw data

MetricClaude Opus 5.5 resultComparison or measurement basis
Artificial Analysis Intelligence Index58 (max)v4.3.2; model page snapshot ranked it #1 / 212
Humanity's Last Exam61.4%Fable 5.1's previous high: 59.1%
SciCode66.9%Fable 5.1's previous high: 63.1%
Terminal-Bench 4.059.6%Tied with GPT-6 Astra xhigh; 11 percentage points above Opus 5
AA-Briefcase v1.11,822 Elo143 above Fable 5.1
GDPval-AA v2.11,846 Elo111 above Fable 5.1; 138 above Opus 5
Intelligence Index output per taskAbout 119k tokensOpus 5 max: about 73k; Fable 5.1 max: about 78k; GPT-6 Astra max: about 27k
Weighted average cost per task$5.98Model page collection snapshot
Total cost of the full index evaluation$8,708.20Model page collection snapshot
Cumulative output for index evaluation260M tokensCumulative measure shown on the model page
API prices$4 input / $20 output / $0.20 cache read, per million tokensPrices listed by Artificial Analysis; not benchmark results for model capability

Conclusions and limitations

This evaluation supports viewing Opus 5.5 max as one of the leading general-purpose models on the Artificial Analysis index at the time, particularly for use cases that prioritize the quality of agentic knowledge work. Terminal-Bench 4.0 shows it close to GPT-6 Astra xhigh, but it produced more than four times as many output tokens per index task as Astra max. Actual costs also depend on request length, cache hits, pricing, and workflow.

These conclusions should be limited to Artificial Analysis's v4.3.2 index and the snapshot visible on 2026-09-23. The ten evaluations do not represent every type of work. AA-Briefcase uses a private test set, and the article does not disclose all inputs, per-task data, or run logs, so external readers cannot fully reproduce the ranking. The reported price changes are pricing information published on the page and should be considered separately from independently measured benchmark performance.

Reproduction notes

First check the Artificial Analysis model page for the Intelligence Index version, evaluation components, model variant, and current dynamic data; then consult its methodology to understand how weighted task costs are calculated. Strictly reproducing the scores would require the version of each benchmark, full prompts, sample counts and weights, fallback behavior, tool configuration, and run logs. The article does not provide all of these. AA-Briefcase uses Stirrup, but its evaluation set is private, so its Elo score cannot be fully rerun based on this note.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Opus 5.5

Use and compare models in Tabbit

Claude Opus 5.5

Related reviews

MediaAnthropic official website2026-09-22

Claude Opus 5.5: Official Benchmarks and Scope

MediaMETR website2026-09-22

METR's Predeployment Evaluation of Claude Opus 5.5

MediaSonarSource official blog2026-09-22

SonarSource: Evaluating Claude Opus 5.5 on Java Code Generation

MediaCodeRabbit official blog2026-09-22

CodeRabbit: Claude Opus 5.5's Recall–Precision Trade-off in Code Review

Claude Opus 5.5

Related prompts

MediaAnthropic Claude Platform Docs

Anthropic’s Prompting Guide for Claude Opus 5.5

MediaAnthropic Claude Platform Docs

Anthropic’s Official Guide to Claude Opus 5.5: New Capabilities and API Configuration

CommunityReddit, GitHub2026-09-22

A Reddit User’s Claude Code Configuration for Claude Opus 5.5

MediaAmazon Bedrock official documentation2026-09-22

Integrating Claude Opus 5.5 with Amazon Bedrock