Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
CommunityClaude Haiku 5.5

Reddit Repeated Test: Claude Code Skill Results Reverse on Haiku 5.5

Original source

Reddit r/ClaudeAI

Authoru/maverickman1111

Tabbit curation1970-01-01

Read original

One-sentence takeaway

This Claude Code skill field report ran the same tasks three times with and without the skill and found that Haiku 5.5's skill benefit could change direction across runs; a single screenshot is not enough to determine whether a skill works.

Use cases

  • Tasks this can help assess: Comparing the same Claude Code skill with and without the skill enabled, and checking whether Haiku 5.5 results remain stable across repeated runs.

  • Tasks not safe to extrapolate to: Generalizing a small number of tasks with one skill to all coding tasks; treating one positive run as a stable improvement; or using this report to rank Haiku 5.5 or judge its overall coding ability.

  • Applicable model versions: Claude Haiku 5.5. The author says these skills were previously tested on Haiku 4.5 and were rerun on Haiku 5.5 in this post.

  • Test environment or client: Claude Code skills. The author tested the 3 skills used in the previous Haiku 4.5 test. The post does not disclose the Claude Code version, complete skill files, model ID, effort, tool set, task inputs, or scoring script.

  • Reasoning tier and parameters: Not stated. The post does not report effort, sampling parameters, tokens, latency, or cost.

Evaluation method

The author reports the following method:

  1. Select the 3 Claude Code skills previously used in the Haiku 4.5 test.

  2. Run exactly the same tasks on Haiku 5.5.

  3. Run each task 3 times with the skill enabled and 3 times without it.

  4. The author says that 4 of the 6 model/task readings changed across identical runs; the post does not show a complete table of the 6 readings.

  5. The author says every result had a dated, hash-verified receipt and that the original post included a link to the complete report. This note relies only on the Reddit body and visible comments and did not open the external report.

Key results

Git workflow skill + Haiku 5.5

The author provides these 3 results:

RunResultAuthor's explanation
Run 1+0.549Clearly helped
Run 2-0.010No clear difference
Run 3+0.235Too few answers to tell

These results show that the same skill and task can show a clear benefit in one run, approach zero in another, and remain inconclusive in a third. The author points out that stopping after Run 1 would produce the conclusion “the skill works very well,” while stopping after Run 2 could produce “the skill is useless.”

Documentation skill

  • On Haiku 4.5, the author says the skill showed clear help in 2 of 3 runs.

  • On Haiku 5.5, the author says all 3 runs were inconclusive.

  • The post does not provide the skill's specific scores, task inputs, or scoring dimensions.

Limitations explicitly stated by the author

The author explicitly says that each skill had only one task and a very small number of answers; this says nothing about Haiku 5.5's overall coding ability and is not a model ranking. In visible comments, the author further explains that 3 repetitions provide only a lower bound on result variability; if the direction can change across 3 runs under the same condition, a single run should not be called a result.

Raw data

  • Run conditions: Same tasks; skill enabled and disabled; 3 repetitions for each condition.

  • Number of skills: 3, from the author's previous Haiku 4.5 test.

  • Run-to-run variation: The author says 4 of the 6 model/task readings changed across identical runs.

  • Git workflow + Haiku 5.5: +0.549, -0.010, +0.235.

  • Documentation skill: Haiku 4.5 showed clear help in 2 of 3 runs; Haiku 5.5 was inconclusive in all 3 runs.

  • Evidence retention: The author says every result had a dated, hash-verified receipt; the body does not show the receipt contents.

  • Unreported fields: Original skill text, task inputs, fixed model IDs, effort, tools, temperature or sampling settings, scoring function, answer count, tokens, latency, cost, run-randomness controls, and statistical intervals.

Conclusions and limits

The report supports the conclusion that skill benefits on Haiku 5.5 vary substantially between runs, and that one successful case or screenshot cannot prove that a skill works. Skill evaluations should retain enabled/disabled comparisons and multiple repetitions at minimum, and report a range or variability rather than a single score.

The report does not support the following conclusions: that a skill is generally effective or ineffective; that Haiku 5.5 is better or worse than Haiku 4.5 overall at coding; Haiku 5.5's performance ranking; or specific changes in the cost, latency, or token usage caused by the skill. The sample contains only a small number of skills and tasks, and the body does not include complete inputs, scoring rules, or run configuration.

Reproduction notes

  1. Fix the Haiku 5.5 model ID, Claude Code version, effort, tools, system prompt, skill version, and task inputs. Save separate configurations with the skill enabled and disabled.

  2. Run the same task for each skill at least 3 times enabled and 3 times disabled. Save every output, timestamp, model-serving information, and hash receipt.

  3. Define scoring dimensions, score range, minimum meaningful gain, and the rules for “inconclusive” before running. Do not change the scoring method after seeing the results.

  4. Record answer count, success rate, error type, tokens, latency, and cost at the same time. If the direction changes across runs, report the range and sample size rather than using the highest score as representative.

  5. Expand to multiple tasks and repositories before judging whether a skill has a stable benefit, and compare Haiku 5.5 with other models using the same tasks, tools, and scoring standard.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Haiku 5.5

Use and compare models in Tabbit

Claude Haiku 5.5

Related reviews

MediaAnthropic2026-10-07

Claude Haiku 5.5 Official Benchmarks: Cost and Capability Positioning for High-Throughput Tasks

MediaArtificial Analysis2026-10-07

Artificial Analysis: Independent Evaluation of Claude Haiku 5.5 on the Intelligence Index and Agent Tasks

CommunityReddit / r/ClaudeCode2026-10-08

Reddit Claude Code Small-Sample Coding-Agent Comparison: Is Haiku 5.5 Medium Good Enough as the Main Model?

CommunityReddit r/ClaudeAI

Reddit User's Claude Code Experience: Haiku 5.5 Context Growth and the 100k Threshold

Claude Haiku 5.5

Related prompts

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Migration Configuration: Switching from Haiku 4.5 to the New API Parameters and Tool Set

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Official Prompting Guide: Effort, Search, and Agent Reliability

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Customer Support Ticket Routing Prompt

CommunityReddit r/ClaudeCode

Reddit Configuration Report: Switching Search Subagents to Haiku 5.5 in Claude Code