Bito ran each task and each model four times on its own service's agent-coding tasks and scored them on a 50-point scale with Opus 5.5. Under each model's default Claude Code settings, Sonnet 5.5 averaged 38.5 versus Sonnet 5's 33.3, with per-session cost about 71% lower and runtime about one-third as long. The default effort differed, however, and the complete tasks, prompts, and scoring implementation are not public.
Tasks this can help assess: Long Claude Code sessions, multi-turn debugging, test-first fixes, design review, vague bug reports, and quick fixes on real service code.
Tasks not safe to extrapolate to: Non-Bito services, other codebases, other agent harnesses, fair model comparisons at the same fixed effort, or actual developer acceptance based only on an Opus 5.5 judge score.
Applicable model versions: Claude Sonnet 5 and Claude Sonnet 5.5. The page also mentions behavioral changes in Sonnet 5 at different times on 2026-09-26, although the API consistently reported Sonnet 5.
Test environment or client: Claude Code; the test used Bito's own agent-coding tasks, fixed prompts, and 4 runs per model for each task. The page does not disclose the complete task code, prompts, answer key, or Claude Code project configuration.
Recommended reasoning tier and parameters: The test retained Claude Code defaults: high effort for Sonnet 5 and medium effort for Sonnet 5.5. Bito recommends that teams use Sonnet 5.5's default directly, but that recommendation is tied to the conditions of this evaluation.
Bito placed the models in the real coding-work benchmark it uses to compare other models. Tasks include short tasks and long multi-turn sessions. A long task runs from locating relevant code, analyzing the failure, planning and implementing a fix, through auditing the fix. The page lists these tasks:
Long session: Slack bot rate limiting;
Long session: tracing a trace ID across two services;
Vague bug report;
Test-first fix;
Design review with author pushback;
Quick fix.
Each model received the same fixed prompts, and each model ran every task four times. The full comparison consumed hundreds of millions of tokens. Each run was scored out of 50 using an answer key that listed the facts required for a correct answer, incorrect claims seen in earlier answers, and the content that the plan or implementation had to complete; every key was checked against the code. The scorer was Claude Opus 5.5, which read the answer, code, and answer key.
Bito kept Claude Code's default settings to simulate most users: high effort for Sonnet 5 and medium effort for Sonnet 5.5. The page says that Claude Code updated from 2.1.283 to 2.1.284 during the Sonnet 5.5 test; the update changed the default effort from high to medium and switched to the shorter system prompt used for Opus. Bito discarded cross-version runs and reran everything on the new version with automatic updates disabled.
| Model and setting | Score (out of 50) | Cost per session | Time per session |
|---|---|---|---|
| Sonnet 5, default (high effort) | 33.3 | $2.65 | 10.7 minutes |
| Sonnet 5.5, default (medium effort) | 38.5 | $0.78 | 3.8 minutes |
Bito reports that Sonnet 5.5 outperformed Sonnet 5 on every task it measured, with session cost about 71% lower and runtime about one-third as long. The page does not publish the complete raw scores and variance for all four runs of every task; it provides only overall and task-level summaries.
| Task | Sonnet 5 | Sonnet 5.5 |
|---|---|---|
| Long session: Slack bot rate limiting | 32.8 | 36.5 |
| Long session: trace ID across two services | 32.6 | 40.1 |
| Vague bug report | 35.2 | 38.0 |
| Test-first fix | 40.8 | 42.0 |
| Design review with pushback | 27.4 | 35.9 |
| Quick fix | 31.1 | 38.4 |
This is the page's largest-improvement scenario: after the agent reviews a pull request, the author replies that “it won't actually have an impact.” The correct response is to return to the code, which shows that the two workflows do conflict.
In 2 of 4 runs, Sonnet 5 accepted the author's claim without checking the code again.
Sonnet 5.5 challenged the claim every time and cited the specific lines showing the conflict.
The score range for this turn was 38–41 for Sonnet 5.5 and 20–35 for Sonnet 5.
The page treats test-first fix as the task with the smallest improvement because it is more mechanical and Sonnet 5 already performed well on it.
For long sessions, Bito recorded:
| Metric | Sonnet 5 | Sonnet 5.5 |
|---|---|---|
| Average thinking tokens / session | About 92,000 | About 25,000 |
| Average model requests / session | 96 | 41 |
| Time per request | About 10 seconds | About 10 seconds |
Bito's explanation is that Sonnet 5.5 takes about the same time per request but makes fewer requests, so the session is faster. This is a workload difference observed in this evaluation, not a latency guarantee for all API workloads.
Based on Claude Code billing observations, Bito says the two models have similar token prices: both cache reads are about $0.20 per million tokens; Sonnet 5.5 output is slightly cheaper and cache writes slightly more expensive. The page attributes the main cost difference to workload rather than token unit price.
In long sessions, Sonnet 5.5 averages about 25,000 thinking tokens and 41 requests, while Sonnet 5 averages about 92,000 thinking tokens and 96 requests. In Bito's measurements, Sonnet 5.5 therefore costs $0.78 per session versus $2.65 for Sonnet 5. These are Bito's Claude Code session snapshots; the page does not disclose complete input, output, cache-hit, retry, or pricing-timestamp data.
Bito also tested Sonnet 5.5 at low effort. Compared with the default medium setting, cost fell by only 9% while the score dropped by nearly 1 point. The author concludes that Sonnet 5.5 is already relatively efficient at medium, leaving limited savings from lowering effort further. The page does not provide the complete score, repetition count, or cost table for this additional run.
| Item | What the Bito page discloses |
|---|---|
| Task source | Real agent-coding tasks on Bito's own service |
| Prompts | All models used the same fixed prompts; the full text is not public |
| Repetitions | 4 per model for each task |
| Maximum score | 50 points |
| Scorer | Claude Opus 5.5, reading the answer, code, and answer key |
| Sonnet 5 default effort | high |
| Sonnet 5.5 default effort | medium |
| Claude Code version | 2.1.283 → 2.1.284; cross-update runs were discarded and everything was rerun with automatic updates disabled on the new version |
| Resource scale | Bito says the comparison consumed hundreds of millions of tokens |
Bito reports run-to-run score noise of about 1–2 points and considers the 5.2-point overall difference in this test clearly above that noise. The author also observed that Sonnet 5 behaved differently on the morning of 2026-09-26 than that afternoon and afterward: the morning version made about half as many requests, used little thinking, and scored about 40, while later runs returned to about 33. The API consistently reported Sonnet 5, and Bito could not prove externally that the model deployment or routing had changed.
Bito's test supports a limited conclusion: on its own service's agent-coding tasks and fixed prompts, Sonnet 5.5 scored higher than Sonnet 5 on every task, with an overall average 5.2 points higher. It also completed sessions with fewer thinking tokens and requests, substantially reducing cost and time. Design review with pushback showed the largest difference, while test-first fix showed the smallest.
This was not a controlled experiment at the same effort: Sonnet 5 used high and Sonnet 5.5 used medium, while the new Claude Code version changed the system prompt and default effort. Opus 5.5 was both the scorer and an external dependency related to model capability. There was no blind human evaluation, public judge prompt, or item-level score, so the 50-point scale cannot be treated as an independent objective measure.
The tasks came from Bito's own service, and the fixed prompts, answer key, and complete code are not public. Cost and time are Claude Code session snapshots affected by model routing, caching, version, retries, tools, and task structure. Sonnet 5's behavioral change on 2026-09-26 also shows that an unchanged model API ID does not guarantee stable behavior over time.
To reproduce Bito's comparison, obtain the same in-house service tasks, fixed prompts, answer keys, code version, and Claude Code version. Run Sonnet 5 at high and Sonnet 5.5 at medium four times per task, then have Opus 5.5 read the answer, code, and key and score them out of 50. Lock the Claude Code version and pass effort explicitly to avoid defaults changing during the test.
Recalculating cost requires recording input, output, thinking, cache-read, and cache-write tokens, model requests, retries, and wall-clock time for each request. To compare the models themselves, run a separate test at the same effort; to compare the default user experience, retain the defaults and report the difference in defaults separately. The Bito page does not publish enough of the tasks, prompts, answer key, or raw outputs for external readers to rerun it completely, so this note classifies it as an independent comparison with a clear method and summarized results rather than a fully reproducible public benchmark.
Claude Sonnet 5.5