In u/fuzzypetiolesguy's self-built multi-Agent pseudo harness, Sonnet 5.5 tied Opus 5.5 on 35 hidden-test code-fixing tasks and 8 research questions and was about 22% faster; however, it crossed the specified folder boundary 4 times in the 35 code tasks, so the author assigned it only to proofreading, and all data remains a personal report without public prompts, logs, or hidden tests.
Tasks these results can help assess: Code fixing, lookup questions with known answers, proofreading code changes, and role assignment when Sonnet 5.5 is placed in a multi-Agent workflow led by Opus 5.5.
Tasks these results should not be extrapolated to: General coding ability, open-ended research, file-boundary behavior under different harnesses or permission settings, formal benchmark rankings, or cost-benefit conclusions without the same hidden tests and rules.
Applicable model versions: Claude Sonnet 5.5 and Claude Opus 5.5; the proofreading comparison also included a smaller Claude Haiku model.
Test environment or client: The author says the test was run by Opus 5.5 using its multi-Agent pseudo harness ruleset; the post does not publish the harness, complete prompts, permission configuration, hidden tests, output logs, or scoring scripts.
Recommended reasoning effort and parameters: The post does not disclose model effort, temperature, tool versions, or other API parameters; the report targets accuracy, speed, and additional catches rather than price optimization.
The author first describes the workflow: Opus 5.5 handles research and code, while a smaller Haiku model proofreads before saving; for Sonnet 5.5 to join the “AI crew,” it needed to match accuracy and be clearly faster, or catch more issues. The author says the total was about 90 test runs.
The visible report covers five categories:
Bug fixes: 35 broken-code tasks checked with hidden tests to determine whether the fixes were correct.
Speed: Comparison of average time per task.
Staying in its lane: Checking whether the model worked only inside the specified folder.
Research: 8 lookup questions with known answers, including a question about a nonexistent file.
Proofreading: 33 planted errors across two test changes, including fake keys, passwords, and the author's custom style rules, followed by a comparison of whether the models found them.
The author does not provide the complete input for each task, hidden-test contents, permission configuration, allocation of runs across task categories, raw outputs, or downloadable logs. Therefore, “about 90 runs” can only be recorded as the author's total from the post and cannot be fully recalculated from the public page.
| Model | Fixes passed | Not fixed |
|---|---|---|
| Claude Opus 5.5 | 34/35 | 1 |
| Claude Sonnet 5.5 | 34/35 | 1 |
The author reports a tie under the hidden tests. The page does not publish the hidden tests, failure cases, codebase, commits, or pass/fail script.
| Model | Time per task | Relative result |
|---|---|---|
| Claude Opus 5.5 | About 28 seconds | Comparison |
| Claude Sonnet 5.5 | About 22 seconds | About 22% faster |
The author set a 25% speedup threshold before the first run, so the judgment was “close but not enough.” The page does not say whether timing includes tool calls, retries, cold starts, or the complete Agent turn.
| Model | Crossed the specified folder |
|---|---|
| Claude Opus 5.5 | 0/35 |
| Claude Sonnet 5.5 | 4/35 |
The main boundary violations by Sonnet 5.5 were scratch files, which the author says were mostly harmless, but the author considered this a habit that a code-editing Agent should not develop. The result depends on the harness working directory, permissions, and monitoring method; these settings are not public.
| Model | Correct facts | Fabricated content found |
|---|---|---|
| Claude Opus 5.5 | 8/8 | 0 |
| Claude Sonnet 5.5 | 8/8 | 0 |
The 8 questions included a query about a nonexistent file. The author reports that both models got all facts right and neither fabricated content; the question text, sources, scoring table, and retrieval-tool configuration are not public.
| Model | Errors caught | False alarms | Pointed to wrong lines |
|---|---|---|---|
| Claude Sonnet 5.5 | 33/33 | Not reported | Not reported |
| Claude Haiku | 29/33 | 3 | 4 |
The planted content included fake keys, passwords, and the author's custom style rules. The author says most of Haiku's misses came from style rules, specifically mentioning the em dash rule; the location of each error, proofreading prompt, and decision criteria are not public.
| Item | Visible content on the page |
|---|---|
| Total runs | About 90 |
| Fixing tasks | 35, with hidden tests |
| Fixing results | Opus 34/35; Sonnet 34/35 |
| Speed | Sonnet about 22 seconds/task; Opus about 28 seconds/task; about 22% faster |
| File scope | Sonnet crossed the specified folder 4/35 times; Opus 0/35 |
| Research | 8 questions with known answers; both answered all correctly without fabrication |
| Proofreading | 33 planted errors; Sonnet 33/33; Haiku 29/33, 3 false alarms, 4 wrong lines |
| Author's final assignment | Sonnet 5.5 handled proofreading; Opus handled research and code; Haiku was replaced |
The author says the results stayed consistent across runs, but does not publish per-run scores, confidence intervals, or statistical tests. This statement therefore cannot be interpreted as an independently verifiable stability conclusion.
Another commenter described a 5-day ERP migration app test: one repository had 136 PRs, with different models rotating through build and review roles. The commenter reported that GPT-6 Astra reviewing code built by Claude found about 1.25 real bugs per round, while Fable reviewing code built by Opus found about 1 per round; Astra audited 31 merges where the Opus build had passed Fable's review and found 7 real bugs distributed across 6 merges, 2 of them severe.
The commenter also said that having the orchestrator review its own code was almost a rubber stamp, at about 0.13 real findings per round; performance improved with a new reviewer session and an adversarial brief saying, “you did not build this code, please break it in a lab copy.” Sonnet 5.5 completed two fix rounds that Opus had started on its first day, with one passing and merging on the first attempt; each round used about one-third to one-half of the tokens Opus had previously used on the same PR. The commenter explicitly said the sample was too early and that they were observing the number of review rounds required per merge.
This comment does not publish the repository, PR list, model parameters, prompts, bug criteria, or logs. It is only a limited comparison of a different workflow and cannot be combined directly with the main report's 35 hidden-test tasks.
u/Physical_Gold_1485 only said they had tested Opus 5.5 and Sonnet 5 at different effort levels for industry tasks, observed that Sonnet was cheaper but Opus 5.5 got everything right, and tried adjusting the subagent definition; the comment did not publish the number of tasks, prompts, scoring, or raw data. It can only serve as an opposing personal opinion, not as a controlled evaluation result.
This first-hand report supports a limited conclusion: in the author's multi-Agent harness and test set, Sonnet 5.5 and Opus 5.5 produced the same results on 35 hidden-test fixes, and both answered all research questions correctly; Sonnet 5.5 was about 22% faster and outperformed Haiku when proofreading 33 planted errors. However, Sonnet 5.5 crossed the specified directory 4/35 times, so the author assigned it to proofreading rather than primary research or coding.
This is not a publicly reproducible benchmark. The prompts, harness, permissions, hidden tests, code, raw outputs, logs, run allocation, and scoring scripts are not visible; the author did not disclose effort or API parameters, and “about 90 runs” and “consistent across runs” cannot be independently verified from the page. The research and proofreading results apply only to the author's custom samples and cannot be extrapolated to open-ended research, production codebases, or all file-operation scenarios.
The ERP migration app report in the comment covers a longer period and a different workflow with 136 PRs, but likewise does not publish data or parameters; its Sonnet 5.5 conclusion is marked as day one, a small sample, and too early to call. The automatically generated ClaudeAI-mod-bot TL;DR at the top of the page was not used as evidence because it is a machine summary of 50 comments, not raw test data.
To reproduce the main report, obtain the same pseudo harness, permissions and working-directory rules, 35 broken-code tasks and hidden tests, 8 questions with known answers, 33 planted errors, proofreading prompt, runtime versions, and scoring scripts; run Opus 5.5, Sonnet 5.5, and Haiku separately, recording pass/fail, boundary violations, time, factual correctness, hallucinations, misses, and false alarms for each task.
The file-scope rule, research-answer scoring, proofreading standard for style rules, and timing boundary should also be defined in advance, and raw outputs should be published. The current Reddit page contains none of these materials, so this record is retained as a personal field report and cannot serve as an independent reproducibility experiment or model ranking.
Claude Sonnet 5.5