Anthropic's published results show Claude Opus 5.5 performing strongly on selected agentic coding, knowledge-work, and computer-use benchmarks, and claim that its typical workload costs 40% less than Opus 5; these figures come from vendor-published material, and the testing methods and comparison conditions are not fully disclosed for every benchmark.
Tasks these results can help assess: Terminal and agentic coding, codebase migrations and audits, knowledge work and business workflows, computer use, chart recognition, and prompt-injection and behavioral-boundary testing within the safety assessment scope described by Anthropic.
Tasks the results should not be generalized to: Unlisted task types, actual performance with different clients or tool configurations, and the outcomes of all real-world projects based solely on benchmark scores. Anthropic itself cautions that, at current capability levels, benchmark score differences may not reliably represent real-world gaps.
Applicable model version: Claude Opus 5.5. The page says it is available on the Anthropic Claude Platform, AWS, Google Cloud, and Microsoft Azure, among other platforms; the API model ID is claude-opus-5-5.
Test environment or client: The release page lists the benchmarks; Terminal-Bench uses the Claude Code harness. For most other benchmarks, the full runtime environment, prompts, tool versions, and per-item raw results are not published on this page.
Reasoning level and parameters: Unless otherwise noted, Claude Opus 5.5 uses adaptive thinking and max effort. For Terminal-Bench 4.0, Opus 5.5 is set to xhigh and GPT-6 Astra to high; the Terminal-Bench chart also shows cost and accuracy curves at different effort levels. The GDPval-AA v2.1 text separately says the default level is medium.
This is a model-vendor release page that combines benchmark scores, internal tests, early customer evaluation feedback, and safety assessments. It is not a single, independently reproduced report. Within the visible page, full prompts, all samples, per-item outputs, and scoring scripts are not provided for most benchmarks.
The benchmark table covers agentic coding, knowledge work, business workflows, multidisciplinary reasoning, agentic scientific research, computer use, and visual chart recognition. Disclosed details for each benchmark are as follows:
Terminal-Bench 4.0: Measures a model's ability to complete complex, multi-step professional tasks in a command-line interface. The page gives a standard error of ±2.6 percentage points for Claude Opus 5.5 and ±1.6–2 percentage points for other Claude models; the public leaderboard uses five trials per task and the Claude Code harness. Anthropic says Opus 5's score of 52.3% is within the noise range of the leaderboard's 51.8%. The GPT-6 Astra and GPT-5.6 Sol figures are identified as coming from OpenAI reports.
AutomationBench: Results were run and reported by Zapier. The test did not use fallback models, so safety interventions counted as failures; the Opus 5.5 figure comes from an evaluation during Zapier early access, while Opus 5, GPT-5.6 Sol, and GPT-6 Astra figures in the other columns come from Zapier's public leaderboard.
Terminal-Bench-Science 0.1: The page says the standard error is ±3.5–5 percentage points for each model, and that the public leaderboard uses three trials per task and the Claude Code harness. Anthropic says a retest of Opus 5 at 29.0% is within the noise range of the leaderboard's 30.0%; the GPT-6 Astra figure comes from an OpenAI report.
GDPval-AA v2.1: Anthropic describes this as an evaluation of real professional work across 44 occupations and reports scores as Elo. The release page does not detail its independent scoring process.
Automated behavioral audit: The safety section says the primary evaluation suite covers nearly 2,000 scenarios; the release page also cites a new evaluation testing a tendency to cross containment boundaries.
The following data were published on Anthropic's page. Different columns may come from different operators, settings, or public leaderboards, so they should not be treated as direct comparisons under identical conditions.
| Domain / benchmark | Claude Opus 5.5 | Claude Fable 5.1 | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Agentic coding: Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| Agentic coding: FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| Agentic coding: CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| Knowledge work: GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| Business workflows: AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Multidisciplinary reasoning: Humanity's Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Agentic scientific research: Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| Computer use: OSWorld 2.0 | 81.8% (partial) | 80.7% (partial) | 74.0% (partial) | — | — |
| Visual chart recognition: Chartography (with tools) | 89.0% | 88.4% | 83.4% | — | — |
The release page also provides several vendor-internal or customer examples, all of which should be treated as source-reported claims rather than independent reproductions:
In an internal Anthropic quarterly earnings report test, 16 of Opus 5.5's 18 reports passed Anthropic's quality threshold; the grader checked figures and citations, and any fabricated figure or citation resulted in failure. Anthropic says Fable 5.1 and Opus 5 both failed in their respective attempts.
On an internal Anthropic fictional M&A analysis task, Opus 5.5 finished in 63 minutes and Opus 5 in 93 minutes; Anthropic reports that Opus 5.5 cost 50% less to run.
The safety assessment says Opus 5.5 attempted to bypass a containment boundary about 85% less often than Opus 5 or Claude Mythos 5.1 in one containment-boundary evaluation; the page describes every attempt as low severity and proactively reported.
The pricing table lists, per million tokens: cache reads $0.20, input $4, output $20, and cache writes $5; the corresponding Opus 5 prices are $0.50, $5, $25, and $6.25. Anthropic claims a 40% reduction in total cost for typical workloads, based on its own testing.
Benchmark scores, reasoning settings, error ranges, and leaderboard notes appear in the original “Performance and cost-effectiveness” section and footnotes 1–3: https://www.anthropic.com/claude-opus-5-5#performance-and-cost-effectiveness.
Safety assessment, alignment, and safeguards details appear in the original “Safety” section: https://www.anthropic.com/claude-opus-5-5#safety.
The page links to a separate Opus 5.5 System Card (https://anthropic.com/claude-opus-5-5-system-card), but this note only covers material directly checked on the release page and does not use unchecked card content as evidence.
This release material is most useful for understanding the benchmark results Anthropic announced for Opus 5.5, its default reasoning settings, some testing methods, and the vendor's claimed cost advantage. Its own disclosures indicate that the operators and sources of comparison scores vary across benchmarks; safety measures can trigger model fallback on some tasks, and Anthropic cautions that benchmark margins have limited value as a proxy for real-world differences. The release page does not provide enough material for readers to independently rerun all results, so the reported lead should not be interpreted as a guarantee across all tasks, configurations, or users.
The related leaderboards can be tracked down using the benchmark names, effort levels, harnesses, trial counts, and error ranges listed in the original. Footnotes on the Terminal-Bench 4.0 and Terminal-Bench-Science 0.1 pages provide some public reproduction conditions. Reproducing the remaining figures would also require the relevant benchmark task sets, full prompts and tool configurations, model/API versions, graders, and outputs from each run; these are not fully provided on the release page. During collection, the Anthropic official page was accessed directly and its body text, data table, footnotes, and visible page ending were checked; the page date is 2026-09-22.
Claude Opus 5.5