Across CodeRabbit's 13 known-defect cases, Sonnet 5.5 with thinking on caught 2 more issues than Sonnet 5 with nearly the same actionable precision; across 44 real open-source PRs, Sonnet 5.5 took about half as long to review on average and produced 24% fewer comments, but large-sample judge scoring is still pending, so comment volume cannot be interpreted directly as quality.
Tasks these results can help assess: Known-defect recall, actionable precision, comment volume, nitpick count, latency, token usage, and Claude API call cost in the CodeRabbit code review pipeline.
Tasks these results should not be extrapolated to: General code generation, other review tools, programming languages not covered, real team comment acceptance rates, or defect recall on the 44 PRs before judge scoring is complete.
Applicable model versions: Claude Sonnet 5.5, Claude Sonnet 5, and Claude Opus 5.5; the Opus 5.5 figures come from CodeRabbit's earlier evaluation of the same Signal cases in September 2026.
Test environment or client: CodeRabbit review pipeline. Signal uses verified issues from 13 real open-source PRs; OSS August uses 44 open-source PRs with 85 known issues. All runs replayed frozen file summaries, walkthroughs, and layer grouping.
Recommended reasoning effort and parameters: Sonnet 5.5 with thinking on uses adaptive thinking, with low, medium, and high effort for the trivial, junior, and senior review cohorts respectively; thinking off uses the same effort ladder with thinking disabled at the low, medium, and high levels supported by Sonnet 5.5. Sonnet 5 uses the same effort ladder; the complete API parameters are not public.
CodeRabbit used two benchmarks:
Signal: 13 difficult real open-source PRs, each with one verified issue that the reviewer should detect. Cases came from Elasticsearch, Puma, vLLM, Cilium, axios, and Next.js; 9 had difficulty 3, 3 had difficulty 4, and 1 had difficulty 5.
OSS August: 44 open-source PRs with 85 known issues, covering logic errors, API misuse, race conditions, null references, and security issues. For this larger test, the page reports only comment volume, latency, tokens, and cost because judge scoring is not complete.
In Signal, an independent judge cast three votes for each comment against the known issue, and only a majority of PASS votes counted. Known issues caught is the share of the 13 cases in which at least one regular actionable comment passed; outside-diff and nitpick comments are excluded. Actionable precision is passed actionable comments divided by all actionable comments; it is not the developer acceptance rate. Reported comments is the number of comments still requiring human reading after validation, deduplication, and filtering.
All runs replayed the same recorded file summaries, walkthroughs, and layer grouping, changing only the review model and thinking setting. The page does not publish the complete PR list, judge prompt, individual comments, or all API parameters.
| Configuration | Known issues caught | Caught including outside-diff | Actionable precision | Reported comments | Average time per review |
|---|---|---|---|---|---|
| Sonnet 5.5, thinking on | 6/13 · 46.2% | 8/13 | 41.2% | 17 | 5:27 |
| Sonnet 5.5, thinking off | 5/13 · 38.5% | 7/13 | 38.5% | 13 | 5:58 |
| Sonnet 5 | 4/13 · 30.8% | 4/13 | 40.0% | 15 | 9:55 |
| Opus 5.5 Standard | 8/13 · 61.5% | 10/13 | 66.7% | 21 | — |
| Opus 5.5 Max | 10/13 · 76.9% | 10/13 | 52.0% | 25 | — |
Sonnet 5.5 with thinking on caught 2 more of the 13 known issues than Sonnet 5, with precision of 41.2% versus 40.0% and 17 versus 15 comments.
The models missed different issues: Sonnet 5.5 found 4 that Sonnet 5 missed, but missed 2 that Sonnet 5 found. Four of the 13 difficult cases defeated every Sonnet configuration.
Sonnet 5.5 with thinking off caught 1 fewer issue than thinking on, had 2.7 percentage points lower precision, produced 4 fewer comments, and took 31 seconds longer on average.
Sonnet 5.5 did not match Opus 5.5: Opus Standard caught 8/13 and Max caught 10/13, both with higher precision.
The two Sonnet 5.5 configurations did not produce identical results: thinking on caught the vLLM config-context bug and a streaming tool-call serialization case that thinking off did not; thinking off caught the Elasticsearch terms-enum case, while thinking on found it only as outside-diff.
| Configuration | Reported comments | Critical / Major / Minor | Nitpicks | Average time | Median | Total time for 44 runs |
|---|---|---|---|---|---|---|
| Sonnet 5.5, thinking on | 111 | 4 / 50 / 57 | 9 | 6:33 | 5:44 | 4 hours 49 minutes |
| Sonnet 5 | 146 | 14 / 87 / 45 | 30 | 13:31 | 13:49 | 9 hours 55 minutes |
CodeRabbit reports that Sonnet 5.5 produced 24% fewer comments than Sonnet 5, about one-third as many nitpicks, and reduced average time from 13:31 to 6:33; Sonnet 5 generated more than 10 minutes of output 4 times, while Sonnet 5.5 did not. Judge scoring is not complete for this set, so comment volume and latency should be understood as workload data rather than a confirmed defect-quality ranking.
On Signal, the two Sonnet 5.5 configurations were:
| Configuration | Known issues caught | Actionable precision | Reported comments | Nitpicks | Average time |
|---|---|---|---|---|---|
| Thinking on | 6/13 | 41.2% | 17 | 2 | 5:27 |
| Thinking off | 5/13 | 38.5% | 13 | 3 | 5:58 |
Thinking on caught 1 more issue, had 2.7 percentage points higher precision, and produced 4 more comments. The core review call averaged about 5,800 output tokens with thinking on and about 2,900 tokens with thinking off; the page says that under the same inputs, thinking on did not make the average complete review slower, and that it should remain enabled by default. This conclusion comes only from these 13 cases and CodeRabbit's configuration.
CodeRabbit calculated Claude model call cost using Anthropic's list prices per million tokens: $2 input, $10 output, $0.20 cache read, and $2.50 cache write; shared small-model summarization and validation calls are excluded from the table below.
| Configuration | Signal, 13 runs | Signal per run | OSS August, 44 runs | OSS per run |
|---|---|---|---|---|
| Sonnet 5.5, thinking on | $6.16 | $0.47 | $20.32 | $0.46 |
| Sonnet 5.5, thinking off | $5.37 | $0.41 | — | — |
| Sonnet 5 | $15.06 | $1.16 | $50.95 | $1.16 |
Across both benchmarks, Sonnet 5.5's Claude call cost was about 40% of Sonnet 5's, saving about 60% per review. Thinking on cost about 15% more than thinking off in exchange for 1 additional hit and slightly higher precision. The cost covers Claude model calls only, not the complete evaluation pipeline.
| Average per core review call | Signal input | Signal output | Signal thinking words | OSS input | OSS output | OSS thinking words |
|---|---|---|---|---|---|---|
| Sonnet 5.5, thinking on | 110.7k | 5.8k | 464 | 87.3k | 5.7k | 523 |
| Sonnet 5.5, thinking off | 110.7k | 2.9k | 0 | — | — | — |
| Sonnet 5 | 247.5k | 21.6k | 2,771 | 191.5k | 23.8k | 3,143 |
Each Signal run contains 22 core review calls, and OSS August contains 84. The page defines input as the gross prompt size of uncached tokens plus cache reads and cache writes, and says that Sonnet 5 reads more than twice as many tokens per review as Sonnet 5.5, writes about four times as many, and uses about six times as many thinking words. Across the complete pipeline, Sonnet 5 used 27% more total tokens than Sonnet 5.5 on Signal and 49% more across the 44 OSS PRs; these totals include the summarization and validation models.
| Metric | Definition on the page |
|---|---|
| Number of cases | 13, each containing 1 verified issue |
| Caught | At least 1 regular actionable comment passed a majority of judge votes |
| Judge | 3 votes per comment; a majority of PASS votes counts |
| Actionable precision | Passed actionable comments / all actionable comments |
| Exclusions | Outside-diff and nitpick comments are excluded from actionable precision; the number of catches including outside-diff is reported separately |
| Run inputs | The same file summaries, walkthrough, and layer grouping in the frozen cassette |
The method notes also point out that the 13-case sample is small and that a single judge call can affect the result. The page gives the Puma case as an example: a finding by Sonnet 5 was accepted, while a similar finding by Sonnet 5.5 was not; 1 of Sonnet 5.5's 7 passed comments received a 2-to-1 pass, which would make precision 35.3% if it were excluded. This shows that 41.2% precision is sensitive to a small number of judgments.
OSS August contains 85 known issues and 44 open-source PRs, but this article publishes only comment volume, severity distribution, nitpicks, latency, tokens, and cost. The page explicitly says judge scoring is pending, so Sonnet 5.5's known-issue recall or precision cannot be calculated from 111 versus 146 comments.
CodeRabbit's Signal test supports a clear but limited conclusion: in its 13 difficult known-defect cases, Sonnet 5.5 with thinking on caught 2/13 more issues than Sonnet 5, with nearly the same precision and shorter average time. Opus 5.5 caught 8/13 or 10/13 in the same cases and had higher precision, so it remains the stronger configuration for difficult reviews.
The 44 real-PR results support fewer comments and shorter review time for Sonnet 5.5 in the CodeRabbit workflow, along with fewer nitpicks; judge scoring for this larger run is not complete. Fewer comments could mean less noise or missed issues, and the current data cannot distinguish between them. Signal precision is also affected by the vote of a three-vote judge, so the page's 35.3% sensitivity example is an important limitation.
The cost results are Claude call costs recalculated from token usage at Anthropic's list prices, using a specific accounting basis for input, output, and cached tokens, and do not include the full evaluation harness cost. The page says the pre-release model used placeholder rates; therefore, the dollar figures cannot be directly extrapolated to other vendors, price versions, cache-hit rates, or actual product configurations.
To reproduce Signal, obtain the same 13 real PRs and verified issues, frozen recorded cassette, CodeRabbit pipeline version, Sonnet 5.5 / Sonnet 5 effort and thinking configuration, tools, and prompts, then have an independent judge cast three votes for each comment. To reproduce OSS August, also obtain the 44 PRs, 85 known issues, and the same validation, deduplication, and filtering rules, and complete judge scoring.
Cost recalculation requires recording each Claude call's uncached input, cache read, cache write, output, thinking tokens, retries, and price version; the page's per-review cost covers Claude model calls only. Sonnet 5.5's actual cost, speed, and quality should also be measured separately on the target team's own PR set; comment count across 44 PRs cannot substitute for quality validation.
Claude Sonnet 5.5