Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Haiku 4.5 · Media / benchmark · Independent measurement

Claude Haiku 4.5: Six Publicly Documented Pieces of Evidence on Claude Haiku 4.5 from BenchLM

As of 2026-08-17, BenchLM shows six source-displayable Haiku 4.5 rows across SWE-bench, VulcanBench, EEBench, JobBench, and FrontierMath; row conditions differ, so the aggregate score cannot replace task-level judgment.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Platform
Six source-displayable BenchLM rows; not one unified harness
Results
SWE-bench, VulcanBench, EEBench, JobBench, and FrontierMath differ in conditions
Configuration
Temperature, effort, repeats, and traces undisclosed
Date
Data through 2026-08-17; dynamic directory needs a fresh check

Key data and applicable tasks

One-sentence takeaway

The BenchLM page shows only six sourced Haiku 4.5 results: 73.3% on SWE-bench Verified, 78.3% on VulcanBench v3, and 16.0% on JobBench, showing that a “fast small model” cannot be evaluated across all tasks using a single coding score.

Test environment

  • Data date: 2026-08-17.

  • Directory status: The page provides 6 source-displayable benchmark rows, a custom public score of 57.06/100, and a rank of #86/218; category weights and model coverage are incomplete.

  • Source types: SWE rows are labeled Provider exact; VulcanBench/EEBench/JobBench rows are labeled Benchmark exact; Math rows come from the Epoch AI leaderboard.

  • Reproduction status: This is a public evidence ledger/aggregation page, not a single run using one unified harness.

Input/configuration

The benchmark-row links on the page point to the Anthropic release page, VulcanBench, EEBench, JobBench, and Epoch AI; prompts, temperature, effort, repeat count, and run traces are not disclosed for each row.

Results data

BenchmarkScorePage evidence/observations
SWE-bench Verified73.3%Provider exact; links to the Anthropic Haiku 4.5 release page
VulcanBench v378.3%Benchmark exact
EEBench V13.5%Benchmark exact; the page shows a clear weakness
JobBench16.0%Benchmark exact; Agentic category
FrontierMath v2 Tiers 1–35.903%Benchmark exact
FrontierMath v2 Tier 42.083%Benchmark exact

Conclusion

Haiku 4.5 has usable scores on some software-engineering and instruction-following tests, but is notably weaker on BenchLM’s Work Agent and FrontierMath entries. It is suited to high-speed, low-cost subtasks rather than serving as the default for every difficult problem.

Limitations

  • BenchLM’s custom score of 57.06/100 and rank of #86/218 are affected by weights, coverage, and the dynamic directory, and cannot be treated as ground truth for general capability.

  • Evidence levels, harnesses, and dates differ across rows; cross-row comparisons should return to the original benchmarks.

  • The public page does not include complete raw inputs or failure traces; strict reproduction requires accessing each source.

  • Haiku’s 73.3% and the Anthropic release page’s relative statements about Sonnet/Haiku cannot be combined directly.

Reproduction steps

  1. Open the original link for each row and lock the benchmark version and task split.

  2. Fix the Haiku 4.5 snapshot, API parameters, tool scaffold, and timeout behavior, then rerun SWE, JobBench, and the math/instruction tasks.

  3. Report success rate, cost, latency, retries, and failure types for each item; do not report the custom total while omitting coverage.

  4. Compare with Sonnet 4.5/4.6 under the same harness to confirm whether routing truly delivers a cost/quality advantage.

What this supports

  • Supports separating strengths and weaknesses by public benchmark row.

What this does not support

  • Does not turn six rows into general capability or directly combine them with Anthropic relative claims.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM · BenchLM · Original publication date 2026-08-17 · Site edit date 2026-09-20

Open original source

Claude Haiku 4.5

Compare Claude Haiku 4.5 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Haiku 4.5: What It Is, Costs, and When to Use It

A sourced guide to Claude Haiku 4.5’s 200K context, $1/$5 API pricing, speed, access routes, lifecycle boundary, and escalation choices.

Related reviews

Claude Haiku 4.5: Anthropic's official Claude Haiku 4.5 release: Speed, cost, coding, and computer useAnthropic’s 2025-10-15 release covers Haiku 4.5 SWE-bench, Augment coding, slide-text, and computer-use cases, but prompts, parameters, samples, and harnesses are incomplete; relative performance is publisher/partner reported.Claude Haiku 4.5: Reddit Users' Real-World Experience with Claude Haiku 4.5 and Its Usage LimitsMultiple Reddit Claude.ai/Claude Code users reported Haiku 4.5 writing, translation, long-text, light-coding, and web-search experiences, but reports conflict and lack a shared prompt, snapshot, tools, or repeats; quotas vary by account and time.Claude Haiku 4.5: Claude Haiku 4.5: Pricing, Context, and Batch Agent ConfigurationTurn Claude Haiku 4.5: Pricing, Context, and Batch Agent Configuration into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.Claude Haiku 4.5: Clear Instructions and Tool Boundaries for Low-Latency Tasks with Claude Haiku 4.5Turn Clear Instructions and Tool Boundaries for Low-Latency Tasks with Claude Haiku 4.5 into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.