This Claude Code skill field report ran the same tasks three times with and without the skill and found that Haiku 5.5's skill benefit could change direction across runs; a single screenshot is not enough to determine whether a skill works.
Tasks this can help assess: Comparing the same Claude Code skill with and without the skill enabled, and checking whether Haiku 5.5 results remain stable across repeated runs.
Tasks not safe to extrapolate to: Generalizing a small number of tasks with one skill to all coding tasks; treating one positive run as a stable improvement; or using this report to rank Haiku 5.5 or judge its overall coding ability.
Applicable model versions: Claude Haiku 5.5. The author says these skills were previously tested on Haiku 4.5 and were rerun on Haiku 5.5 in this post.
Test environment or client: Claude Code skills. The author tested the 3 skills used in the previous Haiku 4.5 test. The post does not disclose the Claude Code version, complete skill files, model ID, effort, tool set, task inputs, or scoring script.
Reasoning tier and parameters: Not stated. The post does not report effort, sampling parameters, tokens, latency, or cost.
The author reports the following method:
Select the 3 Claude Code skills previously used in the Haiku 4.5 test.
Run exactly the same tasks on Haiku 5.5.
Run each task 3 times with the skill enabled and 3 times without it.
The author says that 4 of the 6 model/task readings changed across identical runs; the post does not show a complete table of the 6 readings.
The author says every result had a dated, hash-verified receipt and that the original post included a link to the complete report. This note relies only on the Reddit body and visible comments and did not open the external report.
The author provides these 3 results:
| Run | Result | Author's explanation |
|---|---|---|
| Run 1 | +0.549 | Clearly helped |
| Run 2 | -0.010 | No clear difference |
| Run 3 | +0.235 | Too few answers to tell |
These results show that the same skill and task can show a clear benefit in one run, approach zero in another, and remain inconclusive in a third. The author points out that stopping after Run 1 would produce the conclusion “the skill works very well,” while stopping after Run 2 could produce “the skill is useless.”
On Haiku 4.5, the author says the skill showed clear help in 2 of 3 runs.
On Haiku 5.5, the author says all 3 runs were inconclusive.
The post does not provide the skill's specific scores, task inputs, or scoring dimensions.
The author explicitly says that each skill had only one task and a very small number of answers; this says nothing about Haiku 5.5's overall coding ability and is not a model ranking. In visible comments, the author further explains that 3 repetitions provide only a lower bound on result variability; if the direction can change across 3 runs under the same condition, a single run should not be called a result.
Run conditions: Same tasks; skill enabled and disabled; 3 repetitions for each condition.
Number of skills: 3, from the author's previous Haiku 4.5 test.
Run-to-run variation: The author says 4 of the 6 model/task readings changed across identical runs.
Git workflow + Haiku 5.5: +0.549, -0.010, +0.235.
Documentation skill: Haiku 4.5 showed clear help in 2 of 3 runs; Haiku 5.5 was inconclusive in all 3 runs.
Evidence retention: The author says every result had a dated, hash-verified receipt; the body does not show the receipt contents.
Unreported fields: Original skill text, task inputs, fixed model IDs, effort, tools, temperature or sampling settings, scoring function, answer count, tokens, latency, cost, run-randomness controls, and statistical intervals.
The report supports the conclusion that skill benefits on Haiku 5.5 vary substantially between runs, and that one successful case or screenshot cannot prove that a skill works. Skill evaluations should retain enabled/disabled comparisons and multiple repetitions at minimum, and report a range or variability rather than a single score.
The report does not support the following conclusions: that a skill is generally effective or ineffective; that Haiku 5.5 is better or worse than Haiku 4.5 overall at coding; Haiku 5.5's performance ranking; or specific changes in the cost, latency, or token usage caused by the skill. The sample contains only a small number of skills and tasks, and the body does not include complete inputs, scoring rules, or run configuration.
Fix the Haiku 5.5 model ID, Claude Code version, effort, tools, system prompt, skill version, and task inputs. Save separate configurations with the skill enabled and disabled.
Run the same task for each skill at least 3 times enabled and 3 times disabled. Save every output, timestamp, model-serving information, and hash receipt.
Define scoring dimensions, score range, minimum meaningful gain, and the rules for “inconclusive” before running. Do not change the scoring method after seeing the results.
Record answer count, success rate, error type, tokens, latency, and cost at the same time. If the direction changes across runs, report the range and sample size rather than using the highest score as representative.
Expand to multiple tasks and repositories before judging whether a skill has a stable benefit, and compare Haiku 5.5 with other models using the same tasks, tools, and scoring standard.
Claude Haiku 5.5