Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
CommunityClaude Sonnet 5.5

Reddit: Three-Run Comparison of Claude Sonnet 5.5 and Claude Opus 5.5 with the Same Skills

Original source

Reddit r/ClaudeAI

Authoru/maverickman1111

Source date2026-09-29

Tabbit curation2026-09-29

Read original

One-sentence takeaway

Across 3 Addy Osmani agent-skills, a small task set, and each model's default Claude Code effort, Sonnet 5.5 reached the score Opus 5.5 achieved with a skill in the no-skill Git workflow test. Scores varied substantially across three runs, however, and Opus 5 served as the judge model, so the test does not establish equal general capability between the models.

Use cases

  • Tasks this can help assess: The local impact of skills, output variability, and rough cost differences between Sonnet 5.5 and Opus 5.5 on three process-oriented agent tasks: code review, Git workflow, and documentation/ADR.

  • Tasks not safe to extrapolate to: Complex reasoning, concurrent debugging, hidden cross-module constraints, open-ended research, other skills, or large-scale production tasks. A rubric score from one judge model should not be treated as an objective quality score.

  • Applicable model versions: Claude Sonnet 5.5 and Claude Opus 5.5. The author also mentions historical results from Sonnet 5 and an older Claude Code version in the comments, but the main test compares the two 5.5 models.

  • Test environment or client: Claude Code; 3 skills from Addy Osmani's agent-skills repository: code review, Git workflow, and docs/ADRs. Each model handled the same tasks with the skill enabled and disabled.

  • Recommended reasoning tier and parameters: The original post did not set effort manually; both models used their respective Claude Code defaults. Because the defaults differ, this is not a fair comparison at the same effort level.

Evaluation method

After Sonnet 5.5 launched, the author compared Claude Sonnet 5.5 and Claude Opus 5.5 using 3 skills. Each model ran the same tasks with the skill disabled and enabled, and the complete test suite was repeated 3 times. Opus 5 scored the answers using a rubric. The author provides three Git workflow scores and reports ranges and run-to-run variation for code review and docs/ADRs.

The author explicitly lists these limitations in the post: only 3 skills, a small task set, scoring by another model, and no effort normalization between the two models; each retained its Claude Code default. The author also emphasizes that scores changed even for the same model, task, and settings, and therefore does not trust a single run.

Comments on the post add that the author independently scored each answer three times with the rubric and allowed an “inconclusive” conclusion instead of forcing a ranking. The author also says that in an earlier report, rescoring saved answers with another grader left 50 of 51 scores on the same side of the conclusion. The author did not separate generation variability from scoring variability for the specific 0.83-to-0.71 example in this post.

Key results

Git workflow: three runs for each model

Scores are in run 1, run 2, run 3 order and come from the Opus 5 judge model described in the post.

ModelWithout skillWith skill
Claude Opus 5.50.51 / 0.49 / 0.580.86 / 0.86 / 0.86
Claude Sonnet 5.50.86 / 0.85 / 0.860.86 / 0.86 / 0.86

The author's direct observation is that Opus 5.5 needed the skill to reach Sonnet 5.5's no-skill Git workflow score; for Sonnet 5.5, the skill barely changed the score on this task set. After loading the skills, the author could not distinguish the two 5.5 models from their outputs in all 3 runs across the 3 skills.

Run-to-run variation

  • In one no-skill code-review comparison, Opus 5.5 scored 0.83 / 0.86 / 0.71 across three runs; the author says the previous Claude Code version had scored 0.68 in a similar test the prior week.

  • The post says that 5 of 9 comparisons would change the conclusion depending on which run was examined.

  • In a comment, the author adds that code review with the skill put both 5.5 models around 0.89–0.91; for docs/ADRs, Sonnet 5 scored 0.60–0.70, Sonnet 5.5 0.78–0.85, and Opus 5.5 0.65–0.85, all varying by run. The author says Sonnet 5.5 exceeded Sonnet 5 in 2 of 3 docs-task runs.

Estimated cost

The author estimates the API cost of generating the answers as follows:

ModelCost for the full generation
Claude Sonnet 5.5About $2.30
Claude Opus 5.5About $5.74

The post provides no request-level tokens, price list, cache hits, task count, or cost-calculation script, so these values are estimates from the author's environment only.

Raw data

Main test design

ItemWhat the post discloses
SkillsCode review, Git workflow, and docs/ADRs from Addy Osmani's agent-skills repository
Comparison modelsClaude Sonnet 5.5 and Claude Opus 5.5
ConditionsEach model with the skill enabled and disabled
Repetitions3 per condition
ScorerClaude Opus 5
EffortEach model's Claude Code default; not normalized to the same value
Git workflow scoresOpus 5.5 without skill 0.51 / 0.49 / 0.58; with skill 0.86 / 0.86 / 0.86; Sonnet 5.5 without skill 0.86 / 0.85 / 0.86; with skill 0.86 / 0.86 / 0.86
Cost estimateSonnet 5.5 about $2.30; Opus 5.5 about $5.74

Visible review material

The post says all figures and raw files are in an open-source tool report provided by the author. This note checked and uses only the Reddit post and its visible comments; it did not open the external report link or count material not shown in the post toward the results.

Conclusions and limits

This test supports a local conclusion: on process-oriented skills and a small task set, Sonnet 5.5's no-skill result can reach Opus 5.5's Git workflow score with the same skill. Across the author's three runs, the outputs did not show a clear difference between the two 5.5 models after loading the skill. The results also show that a single run can change the conclusion, making model-generation variability a variable worth measuring during selection.

The test does not show that Sonnet 5.5 and Opus 5.5 have equal general capability. It covers only 3 skills and a small task set; the judge model is Opus 5, the scores come from a rubric rather than a blind human evaluation or publicly recalculable automated script, and the two models use their respective Claude Code default effort, so model capability and reasoning settings are not isolated. The multiple 0.85–0.86 Git workflow scores may also reflect a ceiling in the rubric, as a comment on the post points out.

The cost figures cannot be extrapolated directly to other requests either. The post does not disclose the complete task count, input and output tokens, caching, price version, retries, tool calls, or cost calculation. To use this for model selection, fix the model, effort, tools, and scorer on your own task set, run at least three repetitions, and report generation variability separately from judge-rescoring variability.

Reproduction notes

To reproduce the main test, obtain the same 3 skills, task set, and Claude Code version. Run Sonnet 5.5 and Opus 5.5 with the skill enabled and disabled, repeat each condition 3 times, and use the same rubric and Opus 5 judge. To address the confounders exposed by the post, also normalize effort between the two models, save all raw answers, blind the model names before rescoring, and record input, output, tool calls, retries, caching, and cost.

The post itself does not publish enough of the task list, skill version, rubric, judge prompt, API parameters, or raw outputs for a reader to rerun it completely from the Reddit page alone. This note therefore classifies it as a reproducible test with a clear method and partial results, not a complete independent reproduction.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 5.5

Use and compare models in Tabbit

Claude Sonnet 5.5

Related reviews

MediaAnthropic official website

Claude Sonnet 5.5 Official Capability Benchmarks and Limitations

MediaArtificial Analysis2026-09-28

Artificial Analysis: Independent Evaluation of Claude Sonnet 5.5's Intelligence Index and Agent Tasks

CommunityX / Arena.ai2026-09-30

Arena.ai Code Arena: Real-World WebDev Task Ranking for Claude Sonnet 5.5 High

MediaCodeRabbit official blog2026-09-28

CodeRabbit: Code Review Comparison of Claude Sonnet 5.5, Sonnet 5, and Opus 5.5

Claude Sonnet 5.5

Related prompts

MediaClaude Platform Docs

Anthropic's Official Prompting Guide: Effort, Initiative, and Tool Use in Claude Sonnet 5.5

MediaClaude Platform Docs

Anthropic's Official Migration Guide: Claude Sonnet 5.5 API Configuration and Breaking Changes

MediaClaude Platform Docs2026-09-28

Anthropic's Official Model Overview: Current Claude Sonnet 5.5 Configuration

CommunityGitHub (original file linked from a Reddit r/ClaudeAI post)

Community Configuration: CLAUDE.md Working Rules for Claude Sonnet 5.5