Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Sonnet 5.5

CodeRabbit: Code Review Comparison of Claude Sonnet 5.5, Sonnet 5, and Opus 5.5

Original source

CodeRabbit official blog

AuthorHendrik Krack

Source date2026-09-28

Tabbit curation2026-09-28

Read original

One-sentence takeaway

Across CodeRabbit's 13 known-defect cases, Sonnet 5.5 with thinking on caught 2 more issues than Sonnet 5 with nearly the same actionable precision; across 44 real open-source PRs, Sonnet 5.5 took about half as long to review on average and produced 24% fewer comments, but large-sample judge scoring is still pending, so comment volume cannot be interpreted directly as quality.

Use cases

  • Tasks these results can help assess: Known-defect recall, actionable precision, comment volume, nitpick count, latency, token usage, and Claude API call cost in the CodeRabbit code review pipeline.

  • Tasks these results should not be extrapolated to: General code generation, other review tools, programming languages not covered, real team comment acceptance rates, or defect recall on the 44 PRs before judge scoring is complete.

  • Applicable model versions: Claude Sonnet 5.5, Claude Sonnet 5, and Claude Opus 5.5; the Opus 5.5 figures come from CodeRabbit's earlier evaluation of the same Signal cases in September 2026.

  • Test environment or client: CodeRabbit review pipeline. Signal uses verified issues from 13 real open-source PRs; OSS August uses 44 open-source PRs with 85 known issues. All runs replayed frozen file summaries, walkthroughs, and layer grouping.

  • Recommended reasoning effort and parameters: Sonnet 5.5 with thinking on uses adaptive thinking, with low, medium, and high effort for the trivial, junior, and senior review cohorts respectively; thinking off uses the same effort ladder with thinking disabled at the low, medium, and high levels supported by Sonnet 5.5. Sonnet 5 uses the same effort ladder; the complete API parameters are not public.

Evaluation method

CodeRabbit used two benchmarks:

  1. Signal: 13 difficult real open-source PRs, each with one verified issue that the reviewer should detect. Cases came from Elasticsearch, Puma, vLLM, Cilium, axios, and Next.js; 9 had difficulty 3, 3 had difficulty 4, and 1 had difficulty 5.

  2. OSS August: 44 open-source PRs with 85 known issues, covering logic errors, API misuse, race conditions, null references, and security issues. For this larger test, the page reports only comment volume, latency, tokens, and cost because judge scoring is not complete.

In Signal, an independent judge cast three votes for each comment against the known issue, and only a majority of PASS votes counted. Known issues caught is the share of the 13 cases in which at least one regular actionable comment passed; outside-diff and nitpick comments are excluded. Actionable precision is passed actionable comments divided by all actionable comments; it is not the developer acceptance rate. Reported comments is the number of comments still requiring human reading after validation, deduplication, and filtering.

All runs replayed the same recorded file summaries, walkthroughs, and layer grouping, changing only the review model and thinking setting. The page does not publish the complete PR list, judge prompt, individual comments, or all API parameters.

Key results

Signal: 13 known-issue cases

ConfigurationKnown issues caughtCaught including outside-diffActionable precisionReported commentsAverage time per review
Sonnet 5.5, thinking on6/13 · 46.2%8/1341.2%175:27
Sonnet 5.5, thinking off5/13 · 38.5%7/1338.5%135:58
Sonnet 54/13 · 30.8%4/1340.0%159:55
Opus 5.5 Standard8/13 · 61.5%10/1366.7%21—
Opus 5.5 Max10/13 · 76.9%10/1352.0%25—
  • Sonnet 5.5 with thinking on caught 2 more of the 13 known issues than Sonnet 5, with precision of 41.2% versus 40.0% and 17 versus 15 comments.

  • The models missed different issues: Sonnet 5.5 found 4 that Sonnet 5 missed, but missed 2 that Sonnet 5 found. Four of the 13 difficult cases defeated every Sonnet configuration.

  • Sonnet 5.5 with thinking off caught 1 fewer issue than thinking on, had 2.7 percentage points lower precision, produced 4 fewer comments, and took 31 seconds longer on average.

  • Sonnet 5.5 did not match Opus 5.5: Opus Standard caught 8/13 and Max caught 10/13, both with higher precision.

  • The two Sonnet 5.5 configurations did not produce identical results: thinking on caught the vLLM config-context bug and a streaming tool-call serialization case that thinking off did not; thinking off caught the Elasticsearch terms-enum case, while thinking on found it only as outside-diff.

44 real open-source PRs

ConfigurationReported commentsCritical / Major / MinorNitpicksAverage timeMedianTotal time for 44 runs
Sonnet 5.5, thinking on1114 / 50 / 5796:335:444 hours 49 minutes
Sonnet 514614 / 87 / 453013:3113:499 hours 55 minutes

CodeRabbit reports that Sonnet 5.5 produced 24% fewer comments than Sonnet 5, about one-third as many nitpicks, and reduced average time from 13:31 to 6:33; Sonnet 5 generated more than 10 minutes of output 4 times, while Sonnet 5.5 did not. Judge scoring is not complete for this set, so comment volume and latency should be understood as workload data rather than a confirmed defect-quality ranking.

Thinking on versus off

On Signal, the two Sonnet 5.5 configurations were:

ConfigurationKnown issues caughtActionable precisionReported commentsNitpicksAverage time
Thinking on6/1341.2%1725:27
Thinking off5/1338.5%1335:58

Thinking on caught 1 more issue, had 2.7 percentage points higher precision, and produced 4 more comments. The core review call averaged about 5,800 output tokens with thinking on and about 2,900 tokens with thinking off; the page says that under the same inputs, thinking on did not make the average complete review slower, and that it should remain enabled by default. This conclusion comes only from these 13 cases and CodeRabbit's configuration.

Claude model call cost

CodeRabbit calculated Claude model call cost using Anthropic's list prices per million tokens: $2 input, $10 output, $0.20 cache read, and $2.50 cache write; shared small-model summarization and validation calls are excluded from the table below.

ConfigurationSignal, 13 runsSignal per runOSS August, 44 runsOSS per run
Sonnet 5.5, thinking on$6.16$0.47$20.32$0.46
Sonnet 5.5, thinking off$5.37$0.41——
Sonnet 5$15.06$1.16$50.95$1.16

Across both benchmarks, Sonnet 5.5's Claude call cost was about 40% of Sonnet 5's, saving about 60% per review. Thinking on cost about 15% more than thinking off in exchange for 1 additional hit and slightly higher precision. The cost covers Claude model calls only, not the complete evaluation pipeline.

Tokens and latency

Average per core review callSignal inputSignal outputSignal thinking wordsOSS inputOSS outputOSS thinking words
Sonnet 5.5, thinking on110.7k5.8k46487.3k5.7k523
Sonnet 5.5, thinking off110.7k2.9k0———
Sonnet 5247.5k21.6k2,771191.5k23.8k3,143

Each Signal run contains 22 core review calls, and OSS August contains 84. The page defines input as the gross prompt size of uncached tokens plus cache reads and cache writes, and says that Sonnet 5 reads more than twice as many tokens per review as Sonnet 5.5, writes about four times as many, and uses about six times as many thinking words. Across the complete pipeline, Sonnet 5 used 27% more total tokens than Sonnet 5.5 on Signal and 49% more across the 44 OSS PRs; these totals include the summarization and validation models.

Raw data

Signal scoring definitions

MetricDefinition on the page
Number of cases13, each containing 1 verified issue
CaughtAt least 1 regular actionable comment passed a majority of judge votes
Judge3 votes per comment; a majority of PASS votes counts
Actionable precisionPassed actionable comments / all actionable comments
ExclusionsOutside-diff and nitpick comments are excluded from actionable precision; the number of catches including outside-diff is reported separately
Run inputsThe same file summaries, walkthrough, and layer grouping in the frozen cassette

The method notes also point out that the 13-case sample is small and that a single judge call can affect the result. The page gives the Puma case as an example: a finding by Sonnet 5 was accepted, while a similar finding by Sonnet 5.5 was not; 1 of Sonnet 5.5's 7 passed comments received a 2-to-1 pass, which would make precision 35.3% if it were excluded. This shows that 41.2% precision is sensitive to a small number of judgments.

Boundaries of the 44 PR test

OSS August contains 85 known issues and 44 open-source PRs, but this article publishes only comment volume, severity distribution, nitpicks, latency, tokens, and cost. The page explicitly says judge scoring is pending, so Sonnet 5.5's known-issue recall or precision cannot be calculated from 111 versus 146 comments.

Conclusions and limitations

CodeRabbit's Signal test supports a clear but limited conclusion: in its 13 difficult known-defect cases, Sonnet 5.5 with thinking on caught 2/13 more issues than Sonnet 5, with nearly the same precision and shorter average time. Opus 5.5 caught 8/13 or 10/13 in the same cases and had higher precision, so it remains the stronger configuration for difficult reviews.

The 44 real-PR results support fewer comments and shorter review time for Sonnet 5.5 in the CodeRabbit workflow, along with fewer nitpicks; judge scoring for this larger run is not complete. Fewer comments could mean less noise or missed issues, and the current data cannot distinguish between them. Signal precision is also affected by the vote of a three-vote judge, so the page's 35.3% sensitivity example is an important limitation.

The cost results are Claude call costs recalculated from token usage at Anthropic's list prices, using a specific accounting basis for input, output, and cached tokens, and do not include the full evaluation harness cost. The page says the pre-release model used placeholder rates; therefore, the dollar figures cannot be directly extrapolated to other vendors, price versions, cache-hit rates, or actual product configurations.

Reproduction notes

To reproduce Signal, obtain the same 13 real PRs and verified issues, frozen recorded cassette, CodeRabbit pipeline version, Sonnet 5.5 / Sonnet 5 effort and thinking configuration, tools, and prompts, then have an independent judge cast three votes for each comment. To reproduce OSS August, also obtain the 44 PRs, 85 known issues, and the same validation, deduplication, and filtering rules, and complete judge scoring.

Cost recalculation requires recording each Claude call's uncached input, cache read, cache write, output, thinking tokens, retries, and price version; the page's per-review cost covers Claude model calls only. Sonnet 5.5's actual cost, speed, and quality should also be measured separately on the target team's own PR set; comment count across 44 PRs cannot substitute for quality validation.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 5.5

Use and compare models in Tabbit

Claude Sonnet 5.5

Related reviews

MediaAnthropic official website

Claude Sonnet 5.5 Official Capability Benchmarks and Limitations

MediaArtificial Analysis2026-09-28

Artificial Analysis: Independent Evaluation of Claude Sonnet 5.5's Intelligence Index and Agent Tasks

CommunityX / Arena.ai2026-09-30

Arena.ai Code Arena: Real-World WebDev Task Ranking for Claude Sonnet 5.5 High

MediaBito official blog2026-09-29

Bito: Four-Agent-Coding-Task Comparison of Claude Sonnet 5.5 and Sonnet 5

Claude Sonnet 5.5

Related prompts

MediaClaude Platform Docs

Anthropic's Official Prompting Guide: Effort, Initiative, and Tool Use in Claude Sonnet 5.5

MediaClaude Platform Docs

Anthropic's Official Migration Guide: Claude Sonnet 5.5 API Configuration and Breaking Changes

MediaClaude Platform Docs2026-09-28

Anthropic's Official Model Overview: Current Claude Sonnet 5.5 Configuration

CommunityGitHub (original file linked from a Reddit r/ClaudeAI post)

Community Configuration: CLAUDE.md Working Rules for Claude Sonnet 5.5