Claude Haiku 4.5 · Media / benchmark · Independent measurement
As of 2026-08-17, BenchLM shows six source-displayable Haiku 4.5 rows across SWE-bench, VulcanBench, EEBench, JobBench, and FrontierMath; row conditions differ, so the aggregate score cannot replace task-level judgment.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
The BenchLM page shows only six sourced Haiku 4.5 results: 73.3% on SWE-bench Verified, 78.3% on VulcanBench v3, and 16.0% on JobBench, showing that a “fast small model” cannot be evaluated across all tasks using a single coding score.
Data date: 2026-08-17.
Directory status: The page provides 6 source-displayable benchmark rows, a custom public score of 57.06/100, and a rank of #86/218; category weights and model coverage are incomplete.
Source types: SWE rows are labeled Provider exact; VulcanBench/EEBench/JobBench rows are labeled Benchmark exact; Math rows come from the Epoch AI leaderboard.
Reproduction status: This is a public evidence ledger/aggregation page, not a single run using one unified harness.
The benchmark-row links on the page point to the Anthropic release page, VulcanBench, EEBench, JobBench, and Epoch AI; prompts, temperature, effort, repeat count, and run traces are not disclosed for each row.
| Benchmark | Score | Page evidence/observations |
|---|---|---|
| SWE-bench Verified | 73.3% | Provider exact; links to the Anthropic Haiku 4.5 release page |
| VulcanBench v3 | 78.3% | Benchmark exact |
| EEBench V1 | 3.5% | Benchmark exact; the page shows a clear weakness |
| JobBench | 16.0% | Benchmark exact; Agentic category |
| FrontierMath v2 Tiers 1–3 | 5.903% | Benchmark exact |
| FrontierMath v2 Tier 4 | 2.083% | Benchmark exact |
Haiku 4.5 has usable scores on some software-engineering and instruction-following tests, but is notably weaker on BenchLM’s Work Agent and FrontierMath entries. It is suited to high-speed, low-cost subtasks rather than serving as the default for every difficult problem.
BenchLM’s custom score of 57.06/100 and rank of #86/218 are affected by weights, coverage, and the dynamic directory, and cannot be treated as ground truth for general capability.
Evidence levels, harnesses, and dates differ across rows; cross-row comparisons should return to the original benchmarks.
The public page does not include complete raw inputs or failure traces; strict reproduction requires accessing each source.
Haiku’s 73.3% and the Anthropic release page’s relative statements about Sonnet/Haiku cannot be combined directly.
Open the original link for each row and lock the benchmark version and task split.
Fix the Haiku 4.5 snapshot, API parameters, tool scaffold, and timeout behavior, then rerun SWE, JobBench, and the math/instruction tasks.
Report success rate, cost, latency, retries, and failure types for each item; do not report the custom total while omitting coverage.
Compare with Sonnet 4.5/4.6 under the same harness to confirm whether routing truly delivers a cost/quality advantage.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
BenchLM · BenchLM · Original publication date 2026-08-17 · Site edit date 2026-09-20
Open original sourceClaude Haiku 4.5
Download the Tabbit client to check model access