Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Sonnet 5.5

Bito: Four-Agent-Coding-Task Comparison of Claude Sonnet 5.5 and Sonnet 5

Original source

Bito official blog

AuthorBito team; signed Anand Das in the article

Source date2026-09-29

Tabbit curation2026-09-29

Read original

One-sentence takeaway

Bito ran each task and each model four times on its own service's agent-coding tasks and scored them on a 50-point scale with Opus 5.5. Under each model's default Claude Code settings, Sonnet 5.5 averaged 38.5 versus Sonnet 5's 33.3, with per-session cost about 71% lower and runtime about one-third as long. The default effort differed, however, and the complete tasks, prompts, and scoring implementation are not public.

Use cases

  • Tasks this can help assess: Long Claude Code sessions, multi-turn debugging, test-first fixes, design review, vague bug reports, and quick fixes on real service code.

  • Tasks not safe to extrapolate to: Non-Bito services, other codebases, other agent harnesses, fair model comparisons at the same fixed effort, or actual developer acceptance based only on an Opus 5.5 judge score.

  • Applicable model versions: Claude Sonnet 5 and Claude Sonnet 5.5. The page also mentions behavioral changes in Sonnet 5 at different times on 2026-09-26, although the API consistently reported Sonnet 5.

  • Test environment or client: Claude Code; the test used Bito's own agent-coding tasks, fixed prompts, and 4 runs per model for each task. The page does not disclose the complete task code, prompts, answer key, or Claude Code project configuration.

  • Recommended reasoning tier and parameters: The test retained Claude Code defaults: high effort for Sonnet 5 and medium effort for Sonnet 5.5. Bito recommends that teams use Sonnet 5.5's default directly, but that recommendation is tied to the conditions of this evaluation.

Evaluation method

Bito placed the models in the real coding-work benchmark it uses to compare other models. Tasks include short tasks and long multi-turn sessions. A long task runs from locating relevant code, analyzing the failure, planning and implementing a fix, through auditing the fix. The page lists these tasks:

  • Long session: Slack bot rate limiting;

  • Long session: tracing a trace ID across two services;

  • Vague bug report;

  • Test-first fix;

  • Design review with author pushback;

  • Quick fix.

Each model received the same fixed prompts, and each model ran every task four times. The full comparison consumed hundreds of millions of tokens. Each run was scored out of 50 using an answer key that listed the facts required for a correct answer, incorrect claims seen in earlier answers, and the content that the plan or implementation had to complete; every key was checked against the code. The scorer was Claude Opus 5.5, which read the answer, code, and answer key.

Bito kept Claude Code's default settings to simulate most users: high effort for Sonnet 5 and medium effort for Sonnet 5.5. The page says that Claude Code updated from 2.1.283 to 2.1.284 during the Sonnet 5.5 test; the update changed the default effort from high to medium and switched to the shorter system prompt used for Opus. Bito discarded cross-version runs and reran everything on the new version with automatic updates disabled.

Key results

All-task summary

Model and settingScore (out of 50)Cost per sessionTime per session
Sonnet 5, default (high effort)33.3$2.6510.7 minutes
Sonnet 5.5, default (medium effort)38.5$0.783.8 minutes

Bito reports that Sonnet 5.5 outperformed Sonnet 5 on every task it measured, with session cost about 71% lower and runtime about one-third as long. The page does not publish the complete raw scores and variance for all four runs of every task; it provides only overall and task-level summaries.

Task-level average scores listed on the page

TaskSonnet 5Sonnet 5.5
Long session: Slack bot rate limiting32.836.5
Long session: trace ID across two services32.640.1
Vague bug report35.238.0
Test-first fix40.842.0
Design review with pushback27.435.9
Quick fix31.138.4

Design review with pushback

This is the page's largest-improvement scenario: after the agent reviews a pull request, the author replies that “it won't actually have an impact.” The correct response is to return to the code, which shows that the two workflows do conflict.

  • In 2 of 4 runs, Sonnet 5 accepted the author's claim without checking the code again.

  • Sonnet 5.5 challenged the claim every time and cited the specific lines showing the conflict.

  • The score range for this turn was 38–41 for Sonnet 5.5 and 20–35 for Sonnet 5.

The page treats test-first fix as the task with the smallest improvement because it is more mechanical and Sonnet 5 already performed well on it.

Thinking tokens and request count

For long sessions, Bito recorded:

MetricSonnet 5Sonnet 5.5
Average thinking tokens / sessionAbout 92,000About 25,000
Average model requests / session9641
Time per requestAbout 10 secondsAbout 10 seconds

Bito's explanation is that Sonnet 5.5 takes about the same time per request but makes fewer requests, so the session is faster. This is a workload difference observed in this evaluation, not a latency guarantee for all API workloads.

Token prices and session cost

Based on Claude Code billing observations, Bito says the two models have similar token prices: both cache reads are about $0.20 per million tokens; Sonnet 5.5 output is slightly cheaper and cache writes slightly more expensive. The page attributes the main cost difference to workload rather than token unit price.

In long sessions, Sonnet 5.5 averages about 25,000 thinking tokens and 41 requests, while Sonnet 5 averages about 92,000 thinking tokens and 96 requests. In Bito's measurements, Sonnet 5.5 therefore costs $0.78 per session versus $2.65 for Sonnet 5. These are Bito's Claude Code session snapshots; the page does not disclose complete input, output, cache-hit, retry, or pricing-timestamp data.

Additional low-effort result

Bito also tested Sonnet 5.5 at low effort. Compared with the default medium setting, cost fell by only 9% while the score dropped by nearly 1 point. The author concludes that Sonnet 5.5 is already relatively efficient at medium, leaving limited savings from lowering effort further. The page does not provide the complete score, repetition count, or cost table for this additional run.

Raw data

Scoring and run conditions

ItemWhat the Bito page discloses
Task sourceReal agent-coding tasks on Bito's own service
PromptsAll models used the same fixed prompts; the full text is not public
Repetitions4 per model for each task
Maximum score50 points
ScorerClaude Opus 5.5, reading the answer, code, and answer key
Sonnet 5 default efforthigh
Sonnet 5.5 default effortmedium
Claude Code version2.1.283 → 2.1.284; cross-update runs were discarded and everything was rerun with automatic updates disabled on the new version
Resource scaleBito says the comparison consumed hundreds of millions of tokens

Result changes and run-to-run variability

Bito reports run-to-run score noise of about 1–2 points and considers the 5.2-point overall difference in this test clearly above that noise. The author also observed that Sonnet 5 behaved differently on the morning of 2026-09-26 than that afternoon and afterward: the morning version made about half as many requests, used little thinking, and scored about 40, while later runs returned to about 33. The API consistently reported Sonnet 5, and Bito could not prove externally that the model deployment or routing had changed.

Conclusions and limits

Bito's test supports a limited conclusion: on its own service's agent-coding tasks and fixed prompts, Sonnet 5.5 scored higher than Sonnet 5 on every task, with an overall average 5.2 points higher. It also completed sessions with fewer thinking tokens and requests, substantially reducing cost and time. Design review with pushback showed the largest difference, while test-first fix showed the smallest.

This was not a controlled experiment at the same effort: Sonnet 5 used high and Sonnet 5.5 used medium, while the new Claude Code version changed the system prompt and default effort. Opus 5.5 was both the scorer and an external dependency related to model capability. There was no blind human evaluation, public judge prompt, or item-level score, so the 50-point scale cannot be treated as an independent objective measure.

The tasks came from Bito's own service, and the fixed prompts, answer key, and complete code are not public. Cost and time are Claude Code session snapshots affected by model routing, caching, version, retries, tools, and task structure. Sonnet 5's behavioral change on 2026-09-26 also shows that an unchanged model API ID does not guarantee stable behavior over time.

Reproduction notes

To reproduce Bito's comparison, obtain the same in-house service tasks, fixed prompts, answer keys, code version, and Claude Code version. Run Sonnet 5 at high and Sonnet 5.5 at medium four times per task, then have Opus 5.5 read the answer, code, and key and score them out of 50. Lock the Claude Code version and pass effort explicitly to avoid defaults changing during the test.

Recalculating cost requires recording input, output, thinking, cache-read, and cache-write tokens, model requests, retries, and wall-clock time for each request. To compare the models themselves, run a separate test at the same effort; to compare the default user experience, retain the defaults and report the difference in defaults separately. The Bito page does not publish enough of the tasks, prompts, answer key, or raw outputs for external readers to rerun it completely, so this note classifies it as an independent comparison with a clear method and summarized results rather than a fully reproducible public benchmark.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 5.5

Use and compare models in Tabbit

Claude Sonnet 5.5

Related reviews

MediaAnthropic official website

Claude Sonnet 5.5 Official Capability Benchmarks and Limitations

MediaArtificial Analysis2026-09-28

Artificial Analysis: Independent Evaluation of Claude Sonnet 5.5's Intelligence Index and Agent Tasks

CommunityX / Arena.ai2026-09-30

Arena.ai Code Arena: Real-World WebDev Task Ranking for Claude Sonnet 5.5 High

MediaCodeRabbit official blog2026-09-28

CodeRabbit: Code Review Comparison of Claude Sonnet 5.5, Sonnet 5, and Opus 5.5

Claude Sonnet 5.5

Related prompts

MediaClaude Platform Docs

Anthropic's Official Prompting Guide: Effort, Initiative, and Tool Use in Claude Sonnet 5.5

MediaClaude Platform Docs

Anthropic's Official Migration Guide: Claude Sonnet 5.5 API Configuration and Breaking Changes

MediaClaude Platform Docs2026-09-28

Anthropic's Official Model Overview: Current Claude Sonnet 5.5 Configuration

CommunityGitHub (original file linked from a Reddit r/ClaudeAI post)

Community Configuration: CLAUDE.md Working Rules for Claude Sonnet 5.5