METR's preliminary evaluation finds that Claude Opus 5.5 makes incremental gains over Fable 5.1 across several AI R&D-related tasks. It may modestly increase researcher productivity, but there is no evidence that it can fully automate AI R&D.
Tasks it can help assess: Long-horizon AI R&D tasks, model training and optimization, conceptual argumentation, game-agent programming, and open-ended research and report writing.
Tasks for which findings should not be generalized: General office work and everyday Q&A, untested software engineering tasks, alignment properties, compliance with Anthropic policy thresholds, and whether all AI R&D work can be automated.
Model version covered: Claude Opus 5.5; compared with Fable 5.1.
Test environment or client: API access provided to METR; testing ran for 10 business days. Specific API parameters, model snapshots, and the complete runtime environment were not disclosed.
Reasoning level and parameters: Not specified.
METR describes this work as a preliminary predeployment evaluation. It primarily gathered evidence about the impact of Claude Opus 5.5 on AI R&D, focusing on two questions: how much R&D acceleration the model itself might enable, and whether AI had already significantly accelerated the development of Claude Opus 5.5.
The capability tests used API access and were conducted over 10 business days. They covered five tasks:
Budget NanoGPT Speedrun: A constrained version of the NanoGPT Speedrun competition task, used to assess AI R&D capabilities.
Language Model Conceptual Argumentation (LMCA): A conceptual reasoning dataset that, according to the source, is described in a 2026 study by Cooper et al.
Train a Program: Train a machine learning model to reproduce the behavior of given software.
Gaming Bot: Write a Python program to control a game through a nonvisual API.
Sunlight: Conduct open-ended research and write a corresponding report.
METR also drew on its earlier research into model capability trends, Anthropic's responses to a questionnaire about capabilities and control factors, and an interview with an Anthropic researcher. A separate preliminary assessment of AI-driven acceleration in Anthropic's internal R&D was conducted by an independent METR team. That team had greater access, but shared only its conclusions with the team behind this article, without supporting evidence or reasoning details. METR therefore treated it as an input but did not directly substantiate the report's claims in this article.
METR says its quantitative evaluation shows Claude Opus 5.5 is an incremental improvement over Fable 5.1, rather than a discontinuous leap. Gains appeared on both verifiable tasks (Budget NanoGPT and Gaming Bot) and harder-to-verify tasks (LMCA and Sunlight). The source does not publish task-level scores, effect sizes, or statistical test results.
METR considers Claude Opus 5.5's potential to accelerate AI R&D slightly greater than Fable 5.1's. It may deliver a somewhat larger productivity boost for researchers and automate a limited part of R&D, but is unlikely to fully automate AI R&D.
METR notes that the model still shows qualitative weaknesses on difficult, long-horizon tasks and open-ended reasoning—areas where experts typically have strengths. These include foresight, prediction, building feedback loops, and research judgment or taste. METR says the available evidence does not show major progress in these judgment capabilities over Fable 5.1.
On whether AI significantly accelerated the development of Claude Opus 5.5, METR cites an estimate from another preliminary report: "about 1.5x overall acceleration in capabilities (i.e., 1.5 years' worth of progress in 1 year), with about a 30% probability of reaching 2x acceleration." METR specifically notes that the report did not state the time period to which the estimate applies, so it is unclear whether the estimate corresponds to the development of Claude Opus 5.5.
In its questionnaire responses and interview, Anthropic said Claude Opus 5.5 continues the Anthropic ECI trend at the Mythos level. This is Anthropic's claim, not a result from the quantitative tests presented by METR in this article.
| Item | Disclosed in the source |
|---|---|
| API testing window | 10 business days |
| Number of capability tasks | 5: Budget NanoGPT Speedrun, LMCA, Train a Program, Gaming Bot, and Sunlight |
| Directly compared model | Fable 5.1 |
| Task-level scores, sample size, and number of repetitions | Not disclosed |
| Prompts, scoring rules, reasoning parameters, and complete test configuration | Not disclosed |
| Internal AI R&D acceleration estimate | About 1.5x overall acceleration; about a 30% probability of reaching 2x; applicable time period unclear, and the source report did not provide the supporting evidence details needed for the team behind this article to verify it |
This report is useful as a preliminary assessment of Claude Opus 5.5 on long-horizon tasks related to AI R&D. METR's core judgment is that the model shows incremental gains across tasks over Fable 5.1, but that this is not enough to conclude that AI R&D can now be fully automated. The source does not provide enough task-level data for independent recalculation, so these results should be treated as METR's qualitative evaluation, not as a score leaderboard that readers can verify.
The disclosure about independence should be read in context: METR says the evaluation was conducted under an unpaid AI R&D evaluation agreement. METR drafted the initial summary, and Anthropic had an opportunity to review and edit its text; the final text was included in the Claude Opus 5.5 system card. METR also explicitly says this work was not intended to verify a claim that Anthropic complied with a particular policy threshold, and does not evaluate whether the model has specific alignment properties.
The internal acceleration estimate came from another METR team, and the team behind this article did not receive its supporting evidence. METR describes including this estimate as a trial version of a more comprehensive evaluation process. The estimate's time horizon is unclear, so it cannot be directly attributed to the development cycle of Claude Opus 5.5.
The source does not disclose complete inputs for the five tasks, dataset versions, task counts, sampling methods, run configurations, or per-task results. The evaluation therefore cannot currently be independently recalculated from this page. Reproducing the comparison would require, at a minimum, the same task versions and scoring criteria, explicit model snapshots for Claude Opus 5.5 and Fable 5.1, API parameters, the number of runs for each task, and the raw outputs. The METR page links directly to the Claude Opus 5.5 system card, but this note describes the evaluation method and conclusions based only on the statements on the METR page above.
Claude Opus 5.5