Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Haiku 5.5

Claude Haiku 5.5 Official Benchmarks: Cost and Capability Positioning for High-Throughput Tasks

Original source

Anthropic

AuthorAnthropic

Source date2026-10-07

Tabbit curation2026-10-07

Read original

One-sentence takeaway

Anthropic positions Claude Haiku 5.5 as a high-throughput model for summarization, compression, database queries, classification, real-time customer support, browser operation, and coding subagents. The launch page's official comparisons show clear gains over Haiku 4.5 on some knowledge-work, computer-use, multidisciplinary-reasoning, terminal-coding, and visual-reasoning benchmarks, while complex agentic coding should still favor a larger model.

Use cases

  • Tasks this can help evaluate: High-volume summarization and compression, classification, database queries, browser operation, real-time customer support, short subagent tasks, and cost-sensitive batch knowledge work.

  • Tasks not safe to extrapolate to: Treating launch-page scores as a guarantee for every production task; extrapolating Haiku 5.5 results directly to complex, long-running agentic coding; or making unconditional comparisons across different tool environments or reasoning tiers.

  • Applicable model versions: Claude Haiku 5.5; the page also lists Haiku 4.5, GPT-6 Luna, and Claude Sonnet 5.5 as comparisons.

  • Test environment or client: The Anthropic launch page lists GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 (Main), and Chartography. The page does not disclose the complete runtime environment.

  • Reasoning tier and parameters: Haiku 5.5 supports adjustable effort, and the charts show Low, Med, High, Xhigh, and Max. The complete parameters corresponding to each result are not disclosed.

Evaluation method

This is an official benchmark summary on Anthropic's model launch page, not an independent rerun. The page lists model scores for knowledge work, computer use, multidisciplinary reasoning, agentic coding, and visual reasoning, and specifies conditions for some tests:

  1. OSWorld 2.1 results are marked Offline subset. The page says the benchmark measures an agent's ability to complete long, multi-step tasks on a real computer.

  2. Humanity's Last Exam reports both no tools and with tools results.

  3. FrontierCode 1.1 uses the page's labeled Main subset; the comparison table marks the Sonnet 5.5 result as Xhigh.

  4. The page also provides score and per-run cost curves for OSWorld 2.1, GDPval-AA, and Humanity's Last Exam across different effort tiers. OSWorld's vertical axis is a partial-credit score, not a uniform accuracy metric. The page does not list every chart value point by point.

  5. The page directs readers to the Haiku 5.5 System Card for full evaluation details. This note collects only what the launch page shows directly.

Key results

Official comparison table

Evaluation categoryBenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
Knowledge workGDPval-AA v2.1162073514371840
Knowledge workAA-Briefcase v1.1157861413361824
Computer useOSWorld 2.1 (Offline subset)72.4%15.7%48.9%83.9%
Multidisciplinary reasoningHumanity's Last Exam (no tools)45.9%10.2%—56.9%
Multidisciplinary reasoningHumanity's Last Exam (with tools)57.4%18.7%—64.5%
Agentic codingTerminal-Bench 4.039.2%0.0%16.4%70.6%
Agentic codingFrontierCode 1.1 (Main)46.4%—42.4%52.1% (Xhigh)
Visual reasoningChartography (no tools)46.4%6.4%29.1%61.6%

The page also gives cost and capability curves, but does not provide a complete table of values for every effort tier in the body text. Its description of OSWorld 2.1 is that the evaluation measures an agent's ability to complete long, multi-step tasks on a real computer.

Pricing and official positioning

The Anthropic launch page gives the following prices per million tokens. Haiku 5.5 input, output, and cache-read prices are split into prompts at or below and above 100k tokens:

ItemHaiku 5.5 (≤100k / >100k)Haiku 4.5Sonnet 5.5
Cache reads$0.01 / $0.05$0.10$0.10
Cache writes$0.125 / $0.625$1.25$2.50
Input tokens$0.10 / $0.50$1.00$2.00
Output tokens$0.50 / $2.50$5.00$10.00

The page says Haiku 5.5 is priced 90% below Haiku 4.5 for requests at or below 100k tokens and 50% below it for requests above 100k. Its footnote says about 90% of Haiku 4.5 requests fall in the lower tier; the newer tokenizer also changes token usage per task. Anthropic therefore estimates about 75% lower average running cost. It calls Haiku 5.5 its fastest model at standard model speeds, while Opus Fast Mode may be faster, and recommends Haiku for narrower work such as subagents, compression, and summarization when a larger model handles the primary task.

Raw data

  • Benchmark names: GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 (Main), and Chartography.

  • Haiku 5.5 values: 1620, 1578, 72.4%, 45.9% (no tools), 57.4% (with tools), 39.2%, 46.4%, and 46.4%.

  • Conditions explicitly stated on the page: OSWorld 2.1 is the Offline subset; Humanity's Last Exam is split into no tools / with tools; Chartography is no tools; the Sonnet 5.5 FrontierCode result is marked Xhigh.

  • Cost data: Haiku 5.5 cache reads are $0.01 / $0.05, cache writes are $0.125 / $0.625, input is $0.10 / $0.50, and output is $0.50 / $2.50. All prices are per million tokens, with the values before and after the slash corresponding to prompts at or below and above 100k tokens.

  • Undisclosed fields: The launch page does not provide the sample count for each benchmark, per-question inputs, random seeds, complete system prompts, tool implementations, hardware, token limits, run counts, confidence intervals, or raw logs.

Conclusions and limitations

The launch page supports the conclusion that Haiku 5.5 improved clearly over Haiku 4.5 on this set of official benchmarks listed by Anthropic, and that the company positions it at the intersection of speed, price, and high-frequency narrow tasks. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2%, below Sonnet 5.5 at 70.6% in the same table, so the page itself presents Sonnet 5.5 and Opus 5.5 as better choices for complex agentic coding.

These results are a snapshot from a vendor launch page. The page does not disclose enough information to reproduce the experiments independently, nor can it be used to estimate the target business's accuracy, latency, total cost, or stability. Different effort levels, tool conditions, data subsets, and undisclosed runtime settings can all affect the comparison; production selection still requires rerunning the evaluation with your own inputs, tools, and cost constraints.

Reproduction notes

  1. Fix the model as Claude Haiku 5.5 and record the specific model ID, effort tier, tool configuration, and token budget used by the API or client.

  2. Start with the OSWorld 2.1, GDPval-AA, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 (Main), and Chartography benchmarks listed on the page. Record the data subset and no tools / with tools conditions.

  3. Save complete inputs, scoring scripts, run counts, failure types, input and output tokens, latency, and cost for each benchmark. Do not invent configurations that the page does not provide.

  4. Compare Haiku 5.5 and the comparison models in the same harness, on the same data version, and at the same effort tier. Report results for different tool conditions separately.

  5. Add your own summarization, classification, browser-operation, and subagent tasks for business acceptance. Do not treat scores from the official launch page as a production commitment.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Haiku 5.5

Use and compare models in Tabbit

Claude Haiku 5.5

Related reviews

MediaArtificial Analysis2026-10-07

Artificial Analysis: Independent Evaluation of Claude Haiku 5.5 on the Intelligence Index and Agent Tasks

CommunityReddit / r/ClaudeCode2026-10-08

Reddit Claude Code Small-Sample Coding-Agent Comparison: Is Haiku 5.5 Medium Good Enough as the Main Model?

CommunityReddit r/ClaudeAI

Reddit User's Claude Code Experience: Haiku 5.5 Context Growth and the 100k Threshold

MediaArena.ai Code Arena2026-10-08

Arena.ai WebDev Public Leaderboard: Claude Haiku 5.5 High's Live Ranking

Claude Haiku 5.5

Related prompts

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Migration Configuration: Switching from Haiku 4.5 to the New API Parameters and Tool Set

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Official Prompting Guide: Effort, Search, and Agent Reliability

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Customer Support Ticket Routing Prompt

CommunityReddit r/ClaudeCode

Reddit Configuration Report: Switching Search Subagents to Haiku 5.5 in Claude Code