Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Sonnet 5.5

Claude Sonnet 5.5 Official Capability Benchmarks and Limitations

Original source

Anthropic official website

AuthorAnthropic

Tabbit curation1970-01-01

Read original

One-sentence takeaway

Anthropic's launch page reports that Claude Sonnet 5.5 outperforms Sonnet 5 on selected agentic coding, knowledge work, computer use, and chart recognition benchmarks, with up to a 30% lower per-task cost through fewer task tokens. The page explicitly compares the two models' input, output, and cache-read prices; these are vendor-published results, and the full conditions for some evaluations, complete samples, and raw outputs are not public.

Use cases

  • Tasks these results can help assess: Multi-step professional tasks in the command line, code repair and codebase understanding, well-scoped professional and knowledge work, computer operation, chart recognition, and the safety and alignment evaluations described on Anthropic's page.

  • Tasks these results should not be extrapolated to: Task types not listed, actual performance with other clients or tool configurations, defect rates in real production codebases, or all complex open-ended work inferred from benchmark scores alone. Anthropic explicitly notes that Opus 5.5 remains stronger on complex open-ended work that requires sustained judgment.

  • Applicable model version: Claude Sonnet 5.5; the developer API model name given on the page is claude-sonnet-5-5. The page says the model is available on Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure.

  • Test environments or clients: The page lists Terminal-Bench 4.0, FrontierCode 1.1 (Main), CursorBench 4.0, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, OSWorld 2.1, and Chartography. Except for the benchmark descriptions, tool settings, and footnotes specifically stated on the page, the complete runtime environments are not disclosed there.

  • Reasoning effort and parameters: The page's aggregate table reports cost curves at different effort levels; Claude apps and Claude Code use Medium by default, while Claude Platform uses High by default. This launch page does not give the temperature, random seed, or complete API parameters.

Evaluation method

This is a model-vendor launch page that summarizes public benchmarks, cost and speed comparisons, automated behavior audits, Anthropic internal tests, and early tester feedback. It is not an independently reproducible report under a uniform configuration. The launch page does not provide complete prompts, all samples, per-item outputs, scoring scripts, or run logs for most evaluations.

The page provides the following visible evaluation details:

  • Terminal-Bench 4.0: Measures the model's ability to complete complex, multi-step professional tasks in a command-line interface. The page describes Medium effort in the Claude app as the default setting for Sonnet 5.5 and says the model exceeds Sonnet 5's best result at this setting while costing less than one-tenth as much per task. A footnote says that the Opus 5.5 value in the table uses Xhigh effort; the page also says that Terminal-Bench and OpenAI have not published a GPT-6 Sol result, so the table reports GPT-5.6 Sol.

  • FrontierCode 1.1 (Main): Evaluates whether code changes can be merged without human edits and penalizes changes outside the task scope. A footnote says Sonnet 5.5 scores lower at Max effort than at Xhigh because it runs Claude Code's code-review skill more often; in the two cases Cognition examined, this behavior caused a timeout or produced extra edits outside the requested scope.

  • GDPval-AA v2.1 and AA-Briefcase v1.1: The page says GDPval-AA covers 44 occupations and 9 major industries. A footnote says Artificial Analysis ran both evaluations on a pre-release deployment on Claude Platform that contained a bug that could harm structured-output requests; Anthropic believes that any effect was small and would underestimate Sonnet 5.5, and the bug was later fixed.

  • Humanity's Last Exam: The table marks this as with tools, but the launch page does not provide the complete tool configuration, number of questions, or scoring details.

  • OSWorld 2.1: The table marks this as partial; the launch page does not give the complete task list, interaction traces, or success-determination script.

  • Chartography: The table marks this as no tools; the launch page does not provide the complete chart samples, scoring script, or per-item results.

  • Automated behavior audit: Anthropic says the audit covered about 1,850 scenarios to evaluate alignment, resistance to misuse, and honesty. The page says Sonnet 5.5 exceeded or matched Sonnet 5 on most metrics; in the containment evaluation, it was close to Opus 5.5 and was the model that probed container boundaries least among the models Anthropic tested. The page also states that no evaluation can reliably capture every failure mode.

Key results

The following are results from the table on Anthropic's launch page. Results with effort, tool, or partial/no-tools annotations apply only under the stated conditions.

Domain / benchmarkClaude Sonnet 5.5Claude Sonnet 5Claude Opus 5.5GPT-6 Sol
Agentic coding: Terminal-Bench 4.070.6%10.3%66.4%¹—
Agentic coding: FrontierCode 1.1 (Main)46.2% Max²; 52.1% Xhigh42.4%54.4%49.3% Max; 52.1% Xhigh
Agentic coding: CursorBench 4.055.5%34.1%57.8%—
Knowledge work: GDPval-AA v2.11844144918461487⁴
Knowledge work: AA-Briefcase v1.11811135918221483⁴
Multidisciplinary reasoning: Humanity's Last Exam64.5% (with tools)54.9% (with tools)67.7% (with tools)—
Computer use: OSWorld 2.180.1% (partial)57.0% (partial)81.8% (partial)—
Visual chart recognition: Chartography61.6% (no tools)15.6% (no tools)64.4% (no tools)53.6%⁴ (no tools)

The page also discloses the following conclusions about cost, speed, and testing:

  • Coding and computer use: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 versus 10.3% for Sonnet 5; on CursorBench 4.0, it scores 55.5% versus 34.1%; on OSWorld 2.1, its partial score is 80.1% versus 57.0%. Anthropic says Sonnet 5.5 is the first Sonnet model to defeat Pokémon Red using only screenshot-based operation, but the launch page does not provide the complete inputs or scoring data for this test.

  • Knowledge work: GDPval-AA v2.1 is 1844 for Sonnet 5.5 versus 1449 for Sonnet 5; AA-Briefcase v1.1 is 1811 versus 1359 for Sonnet 5. The page describes GDPval-AA as covering 44 occupations and 9 major industries and says Sonnet 5.5 approaches Opus 5.5.

  • Cost: The launch page lists cache reads at $0.20 per million tokens, cache writes at $2.50, input at $2, and output at $10. It explicitly says Sonnet 5.5 and Sonnet 5 have the same input, output, and cache-read prices; it does not explicitly compare their cache-write prices. Anthropic's tests show up to 30% lower cost per task, but the page does not publish the complete task set, call logs, or statistical intervals.

  • Speed: Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and was its fastest Sonnet model at the time; the page does not provide a common hardware setup, request lengths, concurrency, or latency distribution.

  • Effort and cost: The page explains that higher effort generally means longer working time and higher per-task cost, while scores usually also increase; lower levels are better suited to routine tasks, while higher levels perform longer reasoning and more checks. The launch page does not provide point-by-point values for all charts.

  • Safety and safeguards: Sonnet 5.5 uses network security safeguards similar to Opus 5.5 and falls back to Sonnet 5 on high-risk cybersecurity tasks; its biological safety safeguards are the same as Sonnet 5's. The page says Sonnet 5.5 is the first Sonnet model to include a safety classifier against reasoning extraction and extends preserved thinking. These are Anthropic's launch and safety claims, not an external independent evaluation.

Raw data

Benchmark table from the launch page

MetricSonnet 5.5Sonnet 5Notes
Terminal-Bench 4.070.6%10.3%Described on the page as agentic coding; the table does not annotate the effort level for Sonnet 5.5
FrontierCode 1.1 (Main)46.2% Max; 52.1% Xhigh42.4%The reason Max is lower than Xhigh is given in footnote 2
CursorBench 4.055.5%34.1%Described on the page as agentic coding
GDPval-AA v2.118441449Artificial Analysis pre-release deployment; a structured-output bug may have lowered Sonnet 5.5's score
AA-Briefcase v1.118111359Artificial Analysis pre-release deployment; a structured-output bug may have lowered Sonnet 5.5's score
Humanity's Last Exam64.5% with tools54.9% with toolsThe page does not provide the complete tool configuration
OSWorld 2.180.1% partial57.0% partialThe page does not provide the complete task list
Chartography61.6% no tools15.6% no toolsThe page does not provide the complete chart samples or scoring script

Pricing

Item (per million tokens)Claude Sonnet 5.5Claude Opus 5.5
Cache reads$0.20$0.20
Cache writes$2.50$5
Input tokens$2$4
Output tokens$10$20

Conclusions and limitations

This official material supports a limited conclusion: on the evaluations Anthropic selected and disclosed, Sonnet 5.5 scores higher than Sonnet 5 on coding, knowledge work, computer use, and chart recognition; the listed input, output, and cache-read prices match Sonnet 5. Anthropic's internal tests also report fewer task tokens, up to a 30% lower per-task cost, and more than 30% faster generation. FrontierCode results are affected by effort and the Claude Code code-review skill, so the Max and Xhigh values cannot be treated as a simple comparison under the same settings.

These results come from Anthropic's launch page, and the tools, effort levels, data sources, and test conditions are not fully consistent across benchmarks. GDPval-AA and AA-Briefcase used a pre-release deployment, and the page's footnote discloses a structured-output bug in that deployment. At least the Terminal-Bench result for Opus 5.5 uses Xhigh; the page does not fully label effort for every other result. Anthropic also explicitly says benchmark scores cover only part of capability, and that Opus 5.5 remains stronger on complex, open-ended work requiring sustained judgment. Therefore, these data do not guarantee the same performance for Sonnet 5.5 on any codebase, client, parameter set, or production workflow.

The approximately 1,850 scenarios, fallback mechanism, container-boundary, and reasoning-extraction statements in the safety section are Anthropic's own evaluation and product-safeguard disclosures. They should not be interpreted as proof that all safety failure modes are covered. The page says its cybersecurity and biological safeguards target a narrow set of high-risk requests and do not affect ordinary software development or most life-sciences work.

Reproduction notes

The launch page makes benchmark names, some tool and effort annotations, aggregate scores, pricing, and several footnotes public, but it does not provide all task inputs, sample sizes, complete prompts, model snapshots, random parameters, tool configurations, scoring scripts, per-item outputs, or run logs. Readers can verify the public numbers against the page's tables and footnotes, but cannot independently rerun the complete results from this page alone.

To reproduce a comparable evaluation, at minimum lock the version and task list for each benchmark, the exact model builds for Sonnet 5.5 and Sonnet 5, effort and other API parameters, tools or harnesses, the number of repeated runs, scoring rules, and the pre-release deployment status of GDPval-AA and AA-Briefcase. FrontierCode also requires recording whether Claude Code's code-review skill was triggered and whether a timeout or out-of-scope edit occurred. These details are not all public on the page.

¹ Anthropic page footnote: The Opus 5.5 result for Terminal-Bench 4.0 uses Xhigh effort and represents that model's highest score.

² Anthropic page footnote: FrontierCode penalizes high-quality or helpful changes that fall outside the requested scope; at Max effort, Sonnet 5.5 ran Claude Code's code-review skill more often, and the two cases Anthropic inspected involved a timeout or extra edits, resulting in a lower score.

³ Anthropic page footnote: Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release Sonnet 5.5 deployment on Claude Platform; that deployment contained a bug that could harm structured-output requests and was later fixed.

⁴ Anthropic page footnote: OpenAI fixed a bug that reduced GPT-6 Sol's image-understanding capability; the relevant official scores from Artificial Analysis and Surge AI may not yet reflect that version change.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 5.5

Use and compare models in Tabbit

Claude Sonnet 5.5

Related reviews

MediaArtificial Analysis2026-09-28

Artificial Analysis: Independent Evaluation of Claude Sonnet 5.5's Intelligence Index and Agent Tasks

CommunityX / Arena.ai2026-09-30

Arena.ai Code Arena: Real-World WebDev Task Ranking for Claude Sonnet 5.5 High

MediaCodeRabbit official blog2026-09-28

CodeRabbit: Code Review Comparison of Claude Sonnet 5.5, Sonnet 5, and Opus 5.5

MediaBito official blog2026-09-29

Bito: Four-Agent-Coding-Task Comparison of Claude Sonnet 5.5 and Sonnet 5

Claude Sonnet 5.5

Related prompts

MediaClaude Platform Docs

Anthropic's Official Prompting Guide: Effort, Initiative, and Tool Use in Claude Sonnet 5.5

MediaClaude Platform Docs

Anthropic's Official Migration Guide: Claude Sonnet 5.5 API Configuration and Breaking Changes

MediaClaude Platform Docs2026-09-28

Anthropic's Official Model Overview: Current Claude Sonnet 5.5 Configuration

CommunityGitHub (original file linked from a Reddit r/ClaudeAI post)

Community Configuration: CLAUDE.md Working Rules for Claude Sonnet 5.5