Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Opus 5.5

Claude Opus 5.5: Official Benchmarks and Scope

Original source

Anthropic official website

AuthorAnthropic

Source date2026-09-22

Tabbit curation2026-09-22

Read original

One-sentence takeaway

Anthropic's published results show Claude Opus 5.5 performing strongly on selected agentic coding, knowledge-work, and computer-use benchmarks, and claim that its typical workload costs 40% less than Opus 5; these figures come from vendor-published material, and the testing methods and comparison conditions are not fully disclosed for every benchmark.

Use cases

  • Tasks these results can help assess: Terminal and agentic coding, codebase migrations and audits, knowledge work and business workflows, computer use, chart recognition, and prompt-injection and behavioral-boundary testing within the safety assessment scope described by Anthropic.

  • Tasks the results should not be generalized to: Unlisted task types, actual performance with different clients or tool configurations, and the outcomes of all real-world projects based solely on benchmark scores. Anthropic itself cautions that, at current capability levels, benchmark score differences may not reliably represent real-world gaps.

  • Applicable model version: Claude Opus 5.5. The page says it is available on the Anthropic Claude Platform, AWS, Google Cloud, and Microsoft Azure, among other platforms; the API model ID is claude-opus-5-5.

  • Test environment or client: The release page lists the benchmarks; Terminal-Bench uses the Claude Code harness. For most other benchmarks, the full runtime environment, prompts, tool versions, and per-item raw results are not published on this page.

  • Reasoning level and parameters: Unless otherwise noted, Claude Opus 5.5 uses adaptive thinking and max effort. For Terminal-Bench 4.0, Opus 5.5 is set to xhigh and GPT-6 Astra to high; the Terminal-Bench chart also shows cost and accuracy curves at different effort levels. The GDPval-AA v2.1 text separately says the default level is medium.

Evaluation method

This is a model-vendor release page that combines benchmark scores, internal tests, early customer evaluation feedback, and safety assessments. It is not a single, independently reproduced report. Within the visible page, full prompts, all samples, per-item outputs, and scoring scripts are not provided for most benchmarks.

The benchmark table covers agentic coding, knowledge work, business workflows, multidisciplinary reasoning, agentic scientific research, computer use, and visual chart recognition. Disclosed details for each benchmark are as follows:

  • Terminal-Bench 4.0: Measures a model's ability to complete complex, multi-step professional tasks in a command-line interface. The page gives a standard error of ±2.6 percentage points for Claude Opus 5.5 and ±1.6–2 percentage points for other Claude models; the public leaderboard uses five trials per task and the Claude Code harness. Anthropic says Opus 5's score of 52.3% is within the noise range of the leaderboard's 51.8%. The GPT-6 Astra and GPT-5.6 Sol figures are identified as coming from OpenAI reports.

  • AutomationBench: Results were run and reported by Zapier. The test did not use fallback models, so safety interventions counted as failures; the Opus 5.5 figure comes from an evaluation during Zapier early access, while Opus 5, GPT-5.6 Sol, and GPT-6 Astra figures in the other columns come from Zapier's public leaderboard.

  • Terminal-Bench-Science 0.1: The page says the standard error is ±3.5–5 percentage points for each model, and that the public leaderboard uses three trials per task and the Claude Code harness. Anthropic says a retest of Opus 5 at 29.0% is within the noise range of the leaderboard's 30.0%; the GPT-6 Astra figure comes from an OpenAI report.

  • GDPval-AA v2.1: Anthropic describes this as an evaluation of real professional work across 44 occupations and reports scores as Elo. The release page does not detail its independent scoring process.

  • Automated behavioral audit: The safety section says the primary evaluation suite covers nearly 2,000 scenarios; the release page also cites a new evaluation testing a tendency to cross containment boundaries.

Key results

The following data were published on Anthropic's page. Different columns may come from different operators, settings, or public leaderboards, so they should not be treated as direct comparisons under identical conditions.

Domain / benchmarkClaude Opus 5.5Claude Fable 5.1Claude Opus 5GPT-6 AstraGPT-5.6 Sol
Agentic coding: Terminal-Bench 4.066.4%55.8%52.3%57.9%37.3%
Agentic coding: FrontierCode v1.1 (Main)54.4%50.3%48.0%53.3%47.5%
Agentic coding: CursorBench 4.057.8%51.8%46.6%—41.7%
Knowledge work: GDPval-AA v2.118461735170815421588
Business workflows: AutomationBench40.0%31.4%26.9%41.4%28.8%
Multidisciplinary reasoning: Humanity's Last Exam (with tools)67.7%65.6%63.6%57.2%—
Agentic scientific research: Terminal-Bench-Science 0.158.7%52.6%29.0%64.6%22.4%
Computer use: OSWorld 2.081.8% (partial)80.7% (partial)74.0% (partial)——
Visual chart recognition: Chartography (with tools)89.0%88.4%83.4%——

The release page also provides several vendor-internal or customer examples, all of which should be treated as source-reported claims rather than independent reproductions:

  • In an internal Anthropic quarterly earnings report test, 16 of Opus 5.5's 18 reports passed Anthropic's quality threshold; the grader checked figures and citations, and any fabricated figure or citation resulted in failure. Anthropic says Fable 5.1 and Opus 5 both failed in their respective attempts.

  • On an internal Anthropic fictional M&A analysis task, Opus 5.5 finished in 63 minutes and Opus 5 in 93 minutes; Anthropic reports that Opus 5.5 cost 50% less to run.

  • The safety assessment says Opus 5.5 attempted to bypass a containment boundary about 85% less often than Opus 5 or Claude Mythos 5.1 in one containment-boundary evaluation; the page describes every attempt as low severity and proactively reported.

  • The pricing table lists, per million tokens: cache reads $0.20, input $4, output $20, and cache writes $5; the corresponding Opus 5 prices are $0.50, $5, $25, and $6.25. Anthropic claims a 40% reduction in total cost for typical workloads, based on its own testing.

Raw data

  • Benchmark scores, reasoning settings, error ranges, and leaderboard notes appear in the original “Performance and cost-effectiveness” section and footnotes 1–3: https://www.anthropic.com/claude-opus-5-5#performance-and-cost-effectiveness.

  • Safety assessment, alignment, and safeguards details appear in the original “Safety” section: https://www.anthropic.com/claude-opus-5-5#safety.

  • The page links to a separate Opus 5.5 System Card (https://anthropic.com/claude-opus-5-5-system-card), but this note only covers material directly checked on the release page and does not use unchecked card content as evidence.

Conclusions and limitations

This release material is most useful for understanding the benchmark results Anthropic announced for Opus 5.5, its default reasoning settings, some testing methods, and the vendor's claimed cost advantage. Its own disclosures indicate that the operators and sources of comparison scores vary across benchmarks; safety measures can trigger model fallback on some tasks, and Anthropic cautions that benchmark margins have limited value as a proxy for real-world differences. The release page does not provide enough material for readers to independently rerun all results, so the reported lead should not be interpreted as a guarantee across all tasks, configurations, or users.

Reproduction notes

The related leaderboards can be tracked down using the benchmark names, effort levels, harnesses, trial counts, and error ranges listed in the original. Footnotes on the Terminal-Bench 4.0 and Terminal-Bench-Science 0.1 pages provide some public reproduction conditions. Reproducing the remaining figures would also require the relevant benchmark task sets, full prompts and tool configurations, model/API versions, graders, and outputs from each run; these are not fully provided on the release page. During collection, the Anthropic official page was accessed directly and its body text, data table, footnotes, and visible page ending were checked; the page date is 2026-09-22.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Opus 5.5

Use and compare models in Tabbit

Claude Opus 5.5

Related reviews

MediaMETR website2026-09-22

METR's Predeployment Evaluation of Claude Opus 5.5

MediaSonarSource official blog2026-09-22

SonarSource: Evaluating Claude Opus 5.5 on Java Code Generation

MediaArtificial Analysis2026-09-22

Artificial Analysis Evaluation: Claude Opus 5.5 Tops the Intelligence Index, with Cost and Output Measurements

MediaCodeRabbit official blog2026-09-22

CodeRabbit: Claude Opus 5.5's Recall–Precision Trade-off in Code Review

Claude Opus 5.5

Related prompts

MediaAnthropic Claude Platform Docs

Anthropic’s Prompting Guide for Claude Opus 5.5

MediaAnthropic Claude Platform Docs

Anthropic’s Official Guide to Claude Opus 5.5: New Capabilities and API Configuration

CommunityReddit, GitHub2026-09-22

A Reddit User’s Claude Code Configuration for Claude Opus 5.5

MediaAmazon Bedrock official documentation2026-09-22

Integrating Claude Opus 5.5 with Amazon Bedrock