The BenchLM page shows only six sourced Haiku 4.5 results: 73.3% on SWE-bench Verified, 78.3% on VulcanBench v3, and 16.0% on JobBench, showing that a “fast small model” cannot be evaluated across all tasks using a single coding score.
Data date: 2026-08-17.
Directory status: The page provides 6 source-displayable benchmark rows, a custom public score of 57.06/100, and a rank of #86/218; category weights and model coverage are incomplete.
Source types: SWE rows are labeled Provider exact; VulcanBench/EEBench/JobBench rows are labeled Benchmark exact; Math rows come from the Epoch AI leaderboard.
Reproduction status: This is a public evidence ledger/aggregation page, not a single run using one unified harness.
The benchmark-row links on the page point to the Anthropic release page, VulcanBench, EEBench, JobBench, and Epoch AI; prompts, temperature, effort, repeat count, and run traces are not disclosed for each row.
| Benchmark | Score | Page evidence/observations |
|---|---|---|
| SWE-bench Verified | 73.3% | Provider exact; links to the Anthropic Haiku 4.5 release page |
| VulcanBench v3 | 78.3% | Benchmark exact |
| EEBench V1 | 3.5% | Benchmark exact; the page shows a clear weakness |
| JobBench | 16.0% | Benchmark exact; Agentic category |
| FrontierMath v2 Tiers 1–3 | 5.903% | Benchmark exact |
| FrontierMath v2 Tier 4 | 2.083% | Benchmark exact |
Haiku 4.5 has usable scores on some software-engineering and instruction-following tests, but is notably weaker on BenchLM’s Work Agent and FrontierMath entries. It is suited to high-speed, low-cost subtasks rather than serving as the default for every difficult problem.
BenchLM’s custom score of 57.06/100 and rank of #86/218 are affected by weights, coverage, and the dynamic directory, and cannot be treated as ground truth for general capability.
Evidence levels, harnesses, and dates differ across rows; cross-row comparisons should return to the original benchmarks.
The public page does not include complete raw inputs or failure traces; strict reproduction requires accessing each source.
Haiku’s 73.3% and the Anthropic release page’s relative statements about Sonnet/Haiku cannot be combined directly.
Open the original link for each row and lock the benchmark version and task split.
Fix the Haiku 4.5 snapshot, API parameters, tool scaffold, and timeout behavior, then rerun SWE, JobBench, and the math/instruction tasks.
Report success rate, cost, latency, retries, and failure types for each item; do not report the custom total while omitting coverage.
Compare with Sonnet 4.5/4.6 under the same harness to confirm whether routing truly delivers a cost/quality advantage.
Claude Haiku 4.5