Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
CommunityClaude Haiku 5.5

Reddit Claude Code Small-Sample Coding-Agent Comparison: Is Haiku 5.5 Medium Good Enough as the Main Model?

Original source

Reddit / r/ClaudeCode

Authoru/Escobar747

Source date2026-10-08

Tabbit curation2026-10-08

Read original

One-sentence takeaway

A Claude Code user ran a small comparison on a real, medium-sized mixed-language codebase. Haiku 5.5 Medium reached the same result as Sonnet 5.5 Medium on one cross-language agent task, while Low skipped verification and the higher tier sharply increased request count and cost. This is a starting point for a personal routing experiment, not a general coding ranking.

Use cases

  • Suitable tasks: Everyday coding-agent work in a codebase with clear conventions that requires running tests and making a small number of cross-file changes; especially as a low-cost Medium default before handing high-risk tasks to a larger model.

  • Not suitable for: Extrapolating one codebase, one agent task, and the author's personal judgment to all languages, repository sizes, tool harnesses, or release processes.

  • Applicable model versions: Claude Haiku 5.5 (Low / Medium / High), Claude Sonnet 5.5 Medium, GPT-6 Luna High, and GPT-6.1 Sol Medium; the specific snapshots are not disclosed.

  • Applicable clients, agents, or APIs: Coding agents in a Claude Code context; the post does not disclose the complete client version, system prompt, tool schema, repository URL, or run logs.

  • Recommended reasoning tier and parameters: In this case, try Haiku 5.5 Medium first. Low needs additional verification, while the higher tier brought no quality gain on this agent task. Production routing still needs to be retested on your own task set.

Test environment and input/configuration

  • Codebase: A real, medium-sized mixed-language codebase with a Rust core and a layer of legacy scripting. The author did not publish the repository name, commit SHA, or code snapshot.

  • Comparison models and tiers: Haiku 5.5 Low, Medium, and High; Sonnet 5.5 Medium; GPT-6 Luna High; GPT-6.1 Sol Medium.

  • Task set: Two Rust bug fixes based on vague bug reports; one code review containing 9 planted bugs and 2 decoys; two Rust implementations based on tricky specifications; one judgment task that deliberately misattributed responsibility and contained ambiguous requirements; and one agent task that migrated logic across languages in a real repository based on unwritten conventions.

  • Evaluation preparation: The author says they wrote the answer key and hidden tests before running any model. The post does not publish the answer key, hidden tests, complete prompts, or turn-by-turn traces.

  • Quality criteria: The author's account visibly considers code correctness, hidden-test results, adherence to repository conventions, verification behavior, honest reporting of unverified parts, and whether the model proactively identified design trade-offs. The post does not publish a formal rubric, blind evaluation process, or independent judge.

Results

Task correctness

  • On tasks the author considered “well specified,” every model reached 100%; Haiku Low missed one compilation error in the code review.

  • On the deliberately ambiguous judgment task, GPT-6 Luna persisted with its own interpretation while the other models identified the ambiguity. The post does not provide each model's complete answer or scoring table.

  • The meaningful separation appeared in the cross-language agent task: Haiku Medium and Sonnet Medium both produced correct code, passed all tests, followed repository conventions, explained what they could not verify, and proactively identified a real design trade-off.

Agent-task request count and cost

Model configurationResult summaryAPI requestsEstimated cost
Haiku 5.5 MediumCompleted the cross-language migration task at the same level as Sonnet Medium29About $0.08
Sonnet 5.5 MediumCompleted the cross-language migration task at the same level as Haiku Medium30About $1.68 (at list price)
Haiku 5.5 LowProduced working code, but skipped verification, missed error-handling fallback, and gave the inaccurate explanation that “the build was stuck for an hour” even though the actual runtime did not support that claimNot disclosedNot disclosed
Haiku 5.5 HighCorrect result, but no additional quality gain385About $1.03

The author therefore estimates that Haiku Medium cost about one-twentieth as much as Sonnet Medium on this task. This ratio applies only to the request counts and list prices in that run and should not be treated as a fixed cost multiplier.

Conclusion

This is a small personal comparison with explicit task design and hidden tests. It supports a limited routing judgment: Haiku 5.5 Medium handled one real cross-language agent task in the author's Rust and legacy-scripting codebase at a per-run cost far below Sonnet's. Low had inconsistent verification discipline, while High added many requests without improving the result on this task. Medium can be a candidate default tier for everyday coding, while high-risk release work still needs a stronger model or independent review.

Limitations

  • The sample is very small: one codebase and one agent task that truly separated the models. The 100% result on explicit tasks also suggests that the test set may not have been large enough to distinguish model capability.

  • The codebase, complete prompts, answer key, hidden tests, per-turn tool calls, raw outputs, and evaluation records are not public. External readers cannot strictly reproduce the 29/30/385 requests or their corresponding costs.

  • The results come from one person's judgment, without a blind evaluation, second judge, repeated runs, or confidence intervals. “Same result” and “no quality gain” should be treated as the author's observations.

  • Provider, cache hits, client version, tool permissions, timeouts, and retry policies are not fully recorded for the models, so cost cannot be compared outside those conditions.

  • The Low failure included missed verification and an inaccurate explanation, but the post does not break down failure severity. High's cost was also affected by one anomalous run with 385 requests; neither observation supports a general claim.

  • Experiences from other Reddit commenters are not included in the results. Those comments do not share a common task, configuration, or verifiable artifact and should not be mixed with the original author's controlled small test.

Reproduction steps

  1. Use a shareable medium-sized mixed-language repository at a fixed commit. Prewrite the answer key, hidden tests, and task scoring table; include vague bugs, planted-bug review, specification implementation, ambiguous judgment, and cross-language migration tasks.

  2. Fix the Claude Code version, tool schema, permissions, context, timeouts, retries, provider, model snapshot, and effort. Run Haiku 5.5 Low / Medium / High, Sonnet 5.5 Medium, and the other comparison models separately.

  3. For every run, record the code diff, test results, verification actions, error handling, design trade-offs, request count, input and output tokens, elapsed time, and actual bill. Repeat every task multiple times.

  4. Score with hidden tests and a blind judge separately: give separate scores for functional correctness, repository conventions, verification completeness, factual statements, and design judgment.

  5. Report success rate, failure severity, request count, and cost distribution for each effort tier. Do not extrapolate a single 100% result or a single cost of roughly one-twentieth into a law about the model.

Source excerpt or observation (short excerpt for compliance only)

The author's core observation is that Haiku 5.5 Medium reached the same result as Sonnet Medium on one cross-language agent task while using 29 requests at a cost of about $0.08. This number represents only that small-sample run.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Haiku 5.5

Use and compare models in Tabbit

Claude Haiku 5.5

Related reviews

MediaAnthropic2026-10-07

Claude Haiku 5.5 Official Benchmarks: Cost and Capability Positioning for High-Throughput Tasks

MediaArtificial Analysis2026-10-07

Artificial Analysis: Independent Evaluation of Claude Haiku 5.5 on the Intelligence Index and Agent Tasks

CommunityReddit r/ClaudeAI

Reddit User's Claude Code Experience: Haiku 5.5 Context Growth and the 100k Threshold

MediaArena.ai Code Arena2026-10-08

Arena.ai WebDev Public Leaderboard: Claude Haiku 5.5 High's Live Ranking

Claude Haiku 5.5

Related prompts

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Migration Configuration: Switching from Haiku 4.5 to the New API Parameters and Tool Set

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Official Prompting Guide: Effort, Search, and Agent Reliability

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Customer Support Ticket Routing Prompt

CommunityReddit r/ClaudeCode

Reddit Configuration Report: Switching Search Subagents to Haiku 5.5 in Claude Code