Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Review
MediaClaude Haiku 5.5

Artificial Analysis: Independent Evaluation of Claude Haiku 5.5 on the Intelligence Index and Agent Tasks

Original source

Artificial Analysis

AuthorArtificial Analysis Team

Source date2026-10-07

Tabbit curation2026-10-07

Read original

One-sentence takeaway

Artificial Analysis's independent evaluation gives Claude Haiku 5.5 a score of 43 at max effort, placing it among the leading small models. It substantially outperforms Haiku 4.5 on knowledge work and terminal use, but uses a weighted average of about 162k output tokens per Intelligence Index task, and safety refusals may have depressed its AutomationBench-AA score.

Use cases

  • Suitable tasks: High-volume knowledge work, summarization and compression, narrow agent sub-tasks, terminal operations, and cost-sensitive batch processing; AA-Briefcase and Terminal-Bench 4.0 are the most relevant reference directions.

  • Not suitable for: Fact-intensive tasks without human verification, workflows that require full SaaS automation coverage, or agents with strict per-task token and latency limits.

  • Applicable model versions: Claude Haiku 5.5, using the evaluation deployment in Artificial Analysis's 2026-10-07 article.

  • Applicable clients, agents, or APIs: The Artificial Analysis Intelligence Index and its private AA-Briefcase harness; the article does not disclose complete API calls, tool schemas, or per-task outputs.

  • Recommended reasoning tier and parameters: Use max for maximum capability; if token usage needs to be controlled, compare high with max first. The article reports 38 points for Haiku 5.5 (high) and GPT-6 Luna (max), with about 55k and 50k tokens per task respectively.

Test environment and input/configuration

  • Evaluation scope: The Artificial Analysis Intelligence Index covers knowledge work, agentic terminal use, automation, coding, and factual knowledge tasks.

  • Main evaluations: AA-Briefcase (real knowledge-work tasks), Terminal-Bench 4.0 (agentic terminal use), AA-Omniscience (factual knowledge and hallucinations), and AutomationBench-AA (agentic SaaS workflows).

  • Full index composition: The chart labels Intelligence Index v4.3.2 and lists 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. AA-Briefcase Elo combines rubric pass rate, analytical quality, and presentation quality.

  • Reasoning tiers: The article compares low, medium, high, xhigh, and max; Haiku 5.5 is the first Haiku to support Anthropic effort settings and adaptive thinking.

  • Context and modalities: The article records a 1M-token context window (Haiku 4.5 has 200k), with text and image input and text output.

  • Score caveat: The prerelease AutomationBench-AA test showed over-refusal; Artificial Analysis believes the 35% result may be understated and plans to rerun it after a fix.

  • Pricing caveat: Input/output pricing for requests at or below 100k tokens was $0.10/$0.50 per million tokens, rising to $0.50/$2.50 above 100k. The article says the site did not yet support tiered pricing at the time, so the temporary cost data did not include the higher tier.

Results

Intelligence Index and agent tasks

EvaluationClaude Haiku 5.5Comparison or note
Artificial Analysis Intelligence Index (max)43GLM-5.3 Flash scored 42, Gemini 3.8 Flash 41, and GPT-6 Luna 38; Kimi K3 scored 44, and Claude Sonnet 5.5 (max) scored 56
AA-Briefcase (max)1578 EloHigher than the Kimi K3 and GLM-5.3 results listed in the article; close to Muse Spark 1.3 (max), and within the confidence intervals of GPT-6 Astra (max) and Claude Fable 5.1 (high)
Terminal-Bench 4.033%Haiku 4.5 scored 0%; GLM-5.3 Flash also scored 33%, Gemini 3.8 Flash 20%, and GPT-6 Luna 13%
AA-Omniscience accuracy36%Gemini 3.8 Flash scored 55% and GPT-6 Luna 44%
AA-Omniscience hallucination rate40%Gemini 3.8 Flash scored 55% and GPT-6 Luna 77%; lower is better for this metric
AutomationBench-AA35%GPT-6 Luna, Gemini 3.8 Flash, and GLM-5.3 Flash scored 53%–60%; the article says Haiku's result was affected by over-refusal

Token usage, effort, and pricing

  • Output tokens: Haiku 5.5 (max) uses a weighted average of about 162k output tokens per Intelligence Index task, roughly three times GPT-6 Luna (max, about 50k); the article says this is even higher than Opus 5.5 (max). This is not a fixed amount for every question.

  • Effort gain: Moving from xhigh to max adds 2 points but increases token usage to about 1.8 times as much. Haiku 5.5 (high) scores 38 points at about 55k tokens per task; GPT-6 Luna (max) also scores 38 points at about 50k tokens per task.

  • Tiered pricing: Input/output costs are $0.10/$0.50 at or below 100k tokens and $0.50/$2.50 above 100k. Cache reads are $0.01/$0.05, and 5-minute cache writes are $0.125/$0.625; the values before and after the slash correspond to prompts at or below and above 100k tokens.

Conclusion

Artificial Analysis's results support treating Haiku 5.5 as a high-throughput small model: it reaches 1578 Elo on AA-Briefcase, improves from Haiku 4.5's 0% to 33% on Terminal-Bench 4.0, and scores 43 on the Intelligence Index, ahead of several models in the same tier. Production deployment must also account for the token budget: the weighted average of roughly 162k output tokens per task at max can weaken the cost advantage of a low unit price. Factual-knowledge accuracy is only 36%, and AutomationBench-AA was affected by over-refusal, so complex automation and high-risk knowledge tasks should retain a stronger model or human review.

Limitations

  • This is an independent Artificial Analysis evaluation article, not a fully reproducible experiment with public code, datasets, and complete prompts. The article does not publish the sample count, per-question inputs, temperature, complete system prompts, tool implementations, token limits, run counts, or confidence intervals for each evaluation.

  • AA-Briefcase is a private Artificial Analysis evaluation. External readers cannot reconstruct the same task set from the article alone, and the endpoints of the “confidence intervals” mentioned in the article are not public.

  • AutomationBench-AA's 35% may have been depressed by prerelease safety refusals. The result after the fix should be treated as pending updated data and not used as a final comparison with the other models' current scores.

  • Intelligence Index, AA-Briefcase, Terminal-Bench, AA-Omniscience, and AutomationBench-AA use different tasks, scoring, and tool conditions. Their scores cannot be interpreted as a single unified success rate.

  • The article body gives Sonnet 5.5 (max) a score of 56, but the accompanying Intelligence Index chart peaks at 46 and does not show Sonnet 5.5. The body also says AA-Briefcase overlaps the confidence intervals of GPT-6 Astra and Claude Fable 5.1, while the chart does not show those comparison models. This record retains the body figures and flags the text–chart discrepancy; the charts do not independently verify those comparisons.

  • The article body's 40% AA-Omniscience figure is the hallucination rate. The chart's roughly 60% figure is the complementary non-hallucination rate (1 - hallucination rate).

  • Tiered pricing was not fully reflected in Artificial Analysis's cost charts when the article was published. Recalculate the actual cost above 100k tokens using the provider's live billing rules.

Reproduction steps

  1. Fix the public API snapshot, provider, effort, adaptive-thinking setting, tool permissions, context window, and token limit for Claude Haiku 5.5.

  2. In the same agent harness, run Terminal-Bench 4.0, AA-Briefcase-style knowledge work, AutomationBench-AA-style SaaS workflows, and AA-Omniscience-style factual Q&A. Clearly label approximate reruns where the private datasets are unavailable.

  3. Save the input, tool calls, refusals, output tokens, reasoning tokens, latency, cost, and human corrections for each task. Repeat each task enough times to estimate variance.

  4. Report factual errors, admissions of uncertainty, and safety refusals separately; do not equate a low hallucination rate directly with high factual accuracy.

  5. Compare the capability gain and token increase from high to max, and recalculate actual cost separately for the at-or-below and above-100k-token tiers.

Source excerpt or observation (short excerpt for compliance only)

The article's key observation is that Haiku 5.5 reaches 43 points among small models but consumes a weighted average of about 162k output tokens per Intelligence Index task at max. This conclusion should be cited together with the tiered pricing and prerelease refusal issue.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Haiku 5.5

Use and compare models in Tabbit

Claude Haiku 5.5

Related reviews

MediaAnthropic2026-10-07

Claude Haiku 5.5 Official Benchmarks: Cost and Capability Positioning for High-Throughput Tasks

CommunityReddit / r/ClaudeCode2026-10-08

Reddit Claude Code Small-Sample Coding-Agent Comparison: Is Haiku 5.5 Medium Good Enough as the Main Model?

CommunityReddit r/ClaudeAI

Reddit User's Claude Code Experience: Haiku 5.5 Context Growth and the 100k Threshold

MediaArena.ai Code Arena2026-10-08

Arena.ai WebDev Public Leaderboard: Claude Haiku 5.5 High's Live Ranking

Claude Haiku 5.5

Related prompts

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Migration Configuration: Switching from Haiku 4.5 to the New API Parameters and Tool Set

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Official Prompting Guide: Effort, Search, and Agent Reliability

MediaAnthropic Claude Platform Docs

Claude Haiku 5.5 Customer Support Ticket Routing Prompt

CommunityReddit r/ClaudeCode

Reddit Configuration Report: Switching Search Subagents to Haiku 5.5 in Claude Code