Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaClaude Haiku 4.5

Six Publicly Documented Pieces of Evidence on Claude Haiku 4.5 from BenchLM

Original source

BenchLM

AuthorBenchLM

Source date2026-08-17

Tabbit curation2026-08-19

Read original

One-sentence takeaway

The BenchLM page shows only six sourced Haiku 4.5 results: 73.3% on SWE-bench Verified, 78.3% on VulcanBench v3, and 16.0% on JobBench, showing that a “fast small model” cannot be evaluated across all tasks using a single coding score.

Test environment

  • Data date: 2026-08-17.

  • Directory status: The page provides 6 source-displayable benchmark rows, a custom public score of 57.06/100, and a rank of #86/218; category weights and model coverage are incomplete.

  • Source types: SWE rows are labeled Provider exact; VulcanBench/EEBench/JobBench rows are labeled Benchmark exact; Math rows come from the Epoch AI leaderboard.

  • Reproduction status: This is a public evidence ledger/aggregation page, not a single run using one unified harness.

Input/configuration

The benchmark-row links on the page point to the Anthropic release page, VulcanBench, EEBench, JobBench, and Epoch AI; prompts, temperature, effort, repeat count, and run traces are not disclosed for each row.

Results data

BenchmarkScorePage evidence/observations
SWE-bench Verified73.3%Provider exact; links to the Anthropic Haiku 4.5 release page
VulcanBench v378.3%Benchmark exact
EEBench V13.5%Benchmark exact; the page shows a clear weakness
JobBench16.0%Benchmark exact; Agentic category
FrontierMath v2 Tiers 1–35.903%Benchmark exact
FrontierMath v2 Tier 42.083%Benchmark exact

Conclusion

Haiku 4.5 has usable scores on some software-engineering and instruction-following tests, but is notably weaker on BenchLM’s Work Agent and FrontierMath entries. It is suited to high-speed, low-cost subtasks rather than serving as the default for every difficult problem.

Limitations

  • BenchLM’s custom score of 57.06/100 and rank of #86/218 are affected by weights, coverage, and the dynamic directory, and cannot be treated as ground truth for general capability.

  • Evidence levels, harnesses, and dates differ across rows; cross-row comparisons should return to the original benchmarks.

  • The public page does not include complete raw inputs or failure traces; strict reproduction requires accessing each source.

  • Haiku’s 73.3% and the Anthropic release page’s relative statements about Sonnet/Haiku cannot be combined directly.

Reproduction steps

  1. Open the original link for each row and lock the benchmark version and task split.

  2. Fix the Haiku 4.5 snapshot, API parameters, tool scaffold, and timeout behavior, then rerun SWE, JobBench, and the math/instruction tasks.

  3. Report success rate, cost, latency, retries, and failure types for each item; do not report the custom total while omitting coverage.

  4. Compare with Sonnet 4.5/4.6 under the same harness to confirm whether routing truly delivers a cost/quality advantage.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Haiku 4.5

Use and compare models in Tabbit

Claude Haiku 4.5

Related reviews

MediaAnthropic News / Introducing Claude Haiku 4.52025-10-15

Anthropic's official Claude Haiku 4.5 release: Speed, cost, coding, and computer use

CommunityReddit / r/ClaudeAI2025-10-15

Reddit Users' Real-World Experience with Claude Haiku 4.5 and Its Usage Limits

Claude Haiku 4.5

Related prompts

MediaClaude Platform Docs / Prompting best practices

Clear Instructions and Tool Boundaries for Low-Latency Tasks with Claude Haiku 4.5

MediaClaude Platform Docs / Models overview and Anthropic release notes2025-10-15

Claude Haiku 4.5: Pricing, Context, and Batch Agent Configuration