Across 3 Addy Osmani agent-skills, a small task set, and each model's default Claude Code effort, Sonnet 5.5 reached the score Opus 5.5 achieved with a skill in the no-skill Git workflow test. Scores varied substantially across three runs, however, and Opus 5 served as the judge model, so the test does not establish equal general capability between the models.
Tasks this can help assess: The local impact of skills, output variability, and rough cost differences between Sonnet 5.5 and Opus 5.5 on three process-oriented agent tasks: code review, Git workflow, and documentation/ADR.
Tasks not safe to extrapolate to: Complex reasoning, concurrent debugging, hidden cross-module constraints, open-ended research, other skills, or large-scale production tasks. A rubric score from one judge model should not be treated as an objective quality score.
Applicable model versions: Claude Sonnet 5.5 and Claude Opus 5.5. The author also mentions historical results from Sonnet 5 and an older Claude Code version in the comments, but the main test compares the two 5.5 models.
Test environment or client: Claude Code; 3 skills from Addy Osmani's agent-skills repository: code review, Git workflow, and docs/ADRs. Each model handled the same tasks with the skill enabled and disabled.
Recommended reasoning tier and parameters: The original post did not set effort manually; both models used their respective Claude Code defaults. Because the defaults differ, this is not a fair comparison at the same effort level.
After Sonnet 5.5 launched, the author compared Claude Sonnet 5.5 and Claude Opus 5.5 using 3 skills. Each model ran the same tasks with the skill disabled and enabled, and the complete test suite was repeated 3 times. Opus 5 scored the answers using a rubric. The author provides three Git workflow scores and reports ranges and run-to-run variation for code review and docs/ADRs.
The author explicitly lists these limitations in the post: only 3 skills, a small task set, scoring by another model, and no effort normalization between the two models; each retained its Claude Code default. The author also emphasizes that scores changed even for the same model, task, and settings, and therefore does not trust a single run.
Comments on the post add that the author independently scored each answer three times with the rubric and allowed an “inconclusive” conclusion instead of forcing a ranking. The author also says that in an earlier report, rescoring saved answers with another grader left 50 of 51 scores on the same side of the conclusion. The author did not separate generation variability from scoring variability for the specific 0.83-to-0.71 example in this post.
Scores are in run 1, run 2, run 3 order and come from the Opus 5 judge model described in the post.
| Model | Without skill | With skill |
|---|---|---|
| Claude Opus 5.5 | 0.51 / 0.49 / 0.58 | 0.86 / 0.86 / 0.86 |
| Claude Sonnet 5.5 | 0.86 / 0.85 / 0.86 | 0.86 / 0.86 / 0.86 |
The author's direct observation is that Opus 5.5 needed the skill to reach Sonnet 5.5's no-skill Git workflow score; for Sonnet 5.5, the skill barely changed the score on this task set. After loading the skills, the author could not distinguish the two 5.5 models from their outputs in all 3 runs across the 3 skills.
In one no-skill code-review comparison, Opus 5.5 scored 0.83 / 0.86 / 0.71 across three runs; the author says the previous Claude Code version had scored 0.68 in a similar test the prior week.
The post says that 5 of 9 comparisons would change the conclusion depending on which run was examined.
In a comment, the author adds that code review with the skill put both 5.5 models around 0.89–0.91; for docs/ADRs, Sonnet 5 scored 0.60–0.70, Sonnet 5.5 0.78–0.85, and Opus 5.5 0.65–0.85, all varying by run. The author says Sonnet 5.5 exceeded Sonnet 5 in 2 of 3 docs-task runs.
The author estimates the API cost of generating the answers as follows:
| Model | Cost for the full generation |
|---|---|
| Claude Sonnet 5.5 | About $2.30 |
| Claude Opus 5.5 | About $5.74 |
The post provides no request-level tokens, price list, cache hits, task count, or cost-calculation script, so these values are estimates from the author's environment only.
| Item | What the post discloses |
|---|---|
| Skills | Code review, Git workflow, and docs/ADRs from Addy Osmani's agent-skills repository |
| Comparison models | Claude Sonnet 5.5 and Claude Opus 5.5 |
| Conditions | Each model with the skill enabled and disabled |
| Repetitions | 3 per condition |
| Scorer | Claude Opus 5 |
| Effort | Each model's Claude Code default; not normalized to the same value |
| Git workflow scores | Opus 5.5 without skill 0.51 / 0.49 / 0.58; with skill 0.86 / 0.86 / 0.86; Sonnet 5.5 without skill 0.86 / 0.85 / 0.86; with skill 0.86 / 0.86 / 0.86 |
| Cost estimate | Sonnet 5.5 about $2.30; Opus 5.5 about $5.74 |
The post says all figures and raw files are in an open-source tool report provided by the author. This note checked and uses only the Reddit post and its visible comments; it did not open the external report link or count material not shown in the post toward the results.
This test supports a local conclusion: on process-oriented skills and a small task set, Sonnet 5.5's no-skill result can reach Opus 5.5's Git workflow score with the same skill. Across the author's three runs, the outputs did not show a clear difference between the two 5.5 models after loading the skill. The results also show that a single run can change the conclusion, making model-generation variability a variable worth measuring during selection.
The test does not show that Sonnet 5.5 and Opus 5.5 have equal general capability. It covers only 3 skills and a small task set; the judge model is Opus 5, the scores come from a rubric rather than a blind human evaluation or publicly recalculable automated script, and the two models use their respective Claude Code default effort, so model capability and reasoning settings are not isolated. The multiple 0.85–0.86 Git workflow scores may also reflect a ceiling in the rubric, as a comment on the post points out.
The cost figures cannot be extrapolated directly to other requests either. The post does not disclose the complete task count, input and output tokens, caching, price version, retries, tool calls, or cost calculation. To use this for model selection, fix the model, effort, tools, and scorer on your own task set, run at least three repetitions, and report generation variability separately from judge-rescoring variability.
To reproduce the main test, obtain the same 3 skills, task set, and Claude Code version. Run Sonnet 5.5 and Opus 5.5 with the skill enabled and disabled, repeat each condition 3 times, and use the same rubric and Opus 5 judge. To address the confounders exposed by the post, also normalize effort between the two models, save all raw answers, blind the model names before rescoring, and record input, output, tool calls, retries, caching, and cost.
The post itself does not publish enough of the task list, skill version, rubric, judge prompt, API parameters, or raw outputs for a reader to rerun it completely from the Reddit page alone. This note therefore classifies it as a reproducible test with a clear method and partial results, not a complete independent reproduction.
Claude Sonnet 5.5