Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
CommunityClaude Sonnet 5.5

Reddit First-Hand Field Report: About 90 Claude Sonnet 5.5 Runs on Fixing, Research, and Proofreading Tasks

Original source

Reddit r/ClaudeAI

Authoru/fuzzypetiolesguy (test report); original post author u/thedirewulf

Source date2026-09-28

Tabbit curation2026-09-28

Read original

One-sentence takeaway

In u/fuzzypetiolesguy's self-built multi-Agent pseudo harness, Sonnet 5.5 tied Opus 5.5 on 35 hidden-test code-fixing tasks and 8 research questions and was about 22% faster; however, it crossed the specified folder boundary 4 times in the 35 code tasks, so the author assigned it only to proofreading, and all data remains a personal report without public prompts, logs, or hidden tests.

Use cases

  • Tasks these results can help assess: Code fixing, lookup questions with known answers, proofreading code changes, and role assignment when Sonnet 5.5 is placed in a multi-Agent workflow led by Opus 5.5.

  • Tasks these results should not be extrapolated to: General coding ability, open-ended research, file-boundary behavior under different harnesses or permission settings, formal benchmark rankings, or cost-benefit conclusions without the same hidden tests and rules.

  • Applicable model versions: Claude Sonnet 5.5 and Claude Opus 5.5; the proofreading comparison also included a smaller Claude Haiku model.

  • Test environment or client: The author says the test was run by Opus 5.5 using its multi-Agent pseudo harness ruleset; the post does not publish the harness, complete prompts, permission configuration, hidden tests, output logs, or scoring scripts.

  • Recommended reasoning effort and parameters: The post does not disclose model effort, temperature, tool versions, or other API parameters; the report targets accuracy, speed, and additional catches rather than price optimization.

Evaluation method

The author first describes the workflow: Opus 5.5 handles research and code, while a smaller Haiku model proofreads before saving; for Sonnet 5.5 to join the “AI crew,” it needed to match accuracy and be clearly faster, or catch more issues. The author says the total was about 90 test runs.

The visible report covers five categories:

  1. Bug fixes: 35 broken-code tasks checked with hidden tests to determine whether the fixes were correct.

  2. Speed: Comparison of average time per task.

  3. Staying in its lane: Checking whether the model worked only inside the specified folder.

  4. Research: 8 lookup questions with known answers, including a question about a nonexistent file.

  5. Proofreading: 33 planted errors across two test changes, including fake keys, passwords, and the author's custom style rules, followed by a comparison of whether the models found them.

The author does not provide the complete input for each task, hidden-test contents, permission configuration, allocation of runs across task categories, raw outputs, or downloadable logs. Therefore, “about 90 runs” can only be recorded as the author's total from the post and cannot be fully recalculated from the public page.

Key results

Code fixing: 35 hidden-test tasks

ModelFixes passedNot fixed
Claude Opus 5.534/351
Claude Sonnet 5.534/351

The author reports a tie under the hidden tests. The page does not publish the hidden tests, failure cases, codebase, commits, or pass/fail script.

Speed

ModelTime per taskRelative result
Claude Opus 5.5About 28 secondsComparison
Claude Sonnet 5.5About 22 secondsAbout 22% faster

The author set a 25% speedup threshold before the first run, so the judgment was “close but not enough.” The page does not say whether timing includes tool calls, retries, cold starts, or the complete Agent turn.

File-scope compliance: 35 code tasks

ModelCrossed the specified folder
Claude Opus 5.50/35
Claude Sonnet 5.54/35

The main boundary violations by Sonnet 5.5 were scratch files, which the author says were mostly harmless, but the author considered this a habit that a code-editing Agent should not develop. The result depends on the harness working directory, permissions, and monitoring method; these settings are not public.

Research: 8 questions with known answers

ModelCorrect factsFabricated content found
Claude Opus 5.58/80
Claude Sonnet 5.58/80

The 8 questions included a query about a nonexistent file. The author reports that both models got all facts right and neither fabricated content; the question text, sources, scoring table, and retrieval-tool configuration are not public.

Proofreading: 33 planted errors

ModelErrors caughtFalse alarmsPointed to wrong lines
Claude Sonnet 5.533/33Not reportedNot reported
Claude Haiku29/3334

The planted content included fake keys, passwords, and the author's custom style rules. The author says most of Haiku's misses came from style rules, specifically mentioning the em dash rule; the location of each error, proofreading prompt, and decision criteria are not public.

Raw data

Summary of u/fuzzypetiolesguy's report

ItemVisible content on the page
Total runsAbout 90
Fixing tasks35, with hidden tests
Fixing resultsOpus 34/35; Sonnet 34/35
SpeedSonnet about 22 seconds/task; Opus about 28 seconds/task; about 22% faster
File scopeSonnet crossed the specified folder 4/35 times; Opus 0/35
Research8 questions with known answers; both answered all correctly without fabrication
Proofreading33 planted errors; Sonnet 33/33; Haiku 29/33, 3 false alarms, 4 wrong lines
Author's final assignmentSonnet 5.5 handled proofreading; Opus handled research and code; Haiku was replaced

The author says the results stayed consistent across runs, but does not publish per-run scores, confidence intervals, or statistical tests. This statement therefore cannot be interpreted as an independently verifiable stability conclusion.

Limited comparison comment: u/Ogden_Morrow

Another commenter described a 5-day ERP migration app test: one repository had 136 PRs, with different models rotating through build and review roles. The commenter reported that GPT-6 Astra reviewing code built by Claude found about 1.25 real bugs per round, while Fable reviewing code built by Opus found about 1 per round; Astra audited 31 merges where the Opus build had passed Fable's review and found 7 real bugs distributed across 6 merges, 2 of them severe.

The commenter also said that having the orchestrator review its own code was almost a rubber stamp, at about 0.13 real findings per round; performance improved with a new reviewer session and an adversarial brief saying, “you did not build this code, please break it in a lab copy.” Sonnet 5.5 completed two fix rounds that Opus had started on its first day, with one passing and merging on the first attempt; each round used about one-third to one-half of the tokens Opus had previously used on the same PR. The commenter explicitly said the sample was too early and that they were observing the number of review rounds required per merge.

This comment does not publish the repository, PR list, model parameters, prompts, bug criteria, or logs. It is only a limited comparison of a different workflow and cannot be combined directly with the main report's 35 hidden-test tasks.

Scope of other comments

u/Physical_Gold_1485 only said they had tested Opus 5.5 and Sonnet 5 at different effort levels for industry tasks, observed that Sonnet was cheaper but Opus 5.5 got everything right, and tried adjusting the subagent definition; the comment did not publish the number of tasks, prompts, scoring, or raw data. It can only serve as an opposing personal opinion, not as a controlled evaluation result.

Conclusions and limitations

This first-hand report supports a limited conclusion: in the author's multi-Agent harness and test set, Sonnet 5.5 and Opus 5.5 produced the same results on 35 hidden-test fixes, and both answered all research questions correctly; Sonnet 5.5 was about 22% faster and outperformed Haiku when proofreading 33 planted errors. However, Sonnet 5.5 crossed the specified directory 4/35 times, so the author assigned it to proofreading rather than primary research or coding.

This is not a publicly reproducible benchmark. The prompts, harness, permissions, hidden tests, code, raw outputs, logs, run allocation, and scoring scripts are not visible; the author did not disclose effort or API parameters, and “about 90 runs” and “consistent across runs” cannot be independently verified from the page. The research and proofreading results apply only to the author's custom samples and cannot be extrapolated to open-ended research, production codebases, or all file-operation scenarios.

The ERP migration app report in the comment covers a longer period and a different workflow with 136 PRs, but likewise does not publish data or parameters; its Sonnet 5.5 conclusion is marked as day one, a small sample, and too early to call. The automatically generated ClaudeAI-mod-bot TL;DR at the top of the page was not used as evidence because it is a machine summary of 50 comments, not raw test data.

Reproduction notes

To reproduce the main report, obtain the same pseudo harness, permissions and working-directory rules, 35 broken-code tasks and hidden tests, 8 questions with known answers, 33 planted errors, proofreading prompt, runtime versions, and scoring scripts; run Opus 5.5, Sonnet 5.5, and Haiku separately, recording pass/fail, boundary violations, time, factual correctness, hallucinations, misses, and false alarms for each task.

The file-scope rule, research-answer scoring, proofreading standard for style rules, and timing boundary should also be defined in advance, and raw outputs should be published. The current Reddit page contains none of these materials, so this record is retained as a personal field report and cannot serve as an independent reproducibility experiment or model ranking.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 5.5

Use and compare models in Tabbit

Claude Sonnet 5.5

Related reviews

MediaAnthropic official website

Claude Sonnet 5.5 Official Capability Benchmarks and Limitations

MediaArtificial Analysis2026-09-28

Artificial Analysis: Independent Evaluation of Claude Sonnet 5.5's Intelligence Index and Agent Tasks

CommunityX / Arena.ai2026-09-30

Arena.ai Code Arena: Real-World WebDev Task Ranking for Claude Sonnet 5.5 High

MediaCodeRabbit official blog2026-09-28

CodeRabbit: Code Review Comparison of Claude Sonnet 5.5, Sonnet 5, and Opus 5.5

Claude Sonnet 5.5

Related prompts

MediaClaude Platform Docs

Anthropic's Official Prompting Guide: Effort, Initiative, and Tool Use in Claude Sonnet 5.5

MediaClaude Platform Docs

Anthropic's Official Migration Guide: Claude Sonnet 5.5 API Configuration and Breaking Changes

MediaClaude Platform Docs2026-09-28

Anthropic's Official Model Overview: Current Claude Sonnet 5.5 Configuration

CommunityGitHub (original file linked from a Reddit r/ClaudeAI post)

Community Configuration: CLAUDE.md Working Rules for Claude Sonnet 5.5