Anthropic positions Claude Haiku 5.5 as a high-throughput model for summarization, compression, database queries, classification, real-time customer support, browser operation, and coding subagents. The launch page's official comparisons show clear gains over Haiku 4.5 on some knowledge-work, computer-use, multidisciplinary-reasoning, terminal-coding, and visual-reasoning benchmarks, while complex agentic coding should still favor a larger model.
Tasks this can help evaluate: High-volume summarization and compression, classification, database queries, browser operation, real-time customer support, short subagent tasks, and cost-sensitive batch knowledge work.
Tasks not safe to extrapolate to: Treating launch-page scores as a guarantee for every production task; extrapolating Haiku 5.5 results directly to complex, long-running agentic coding; or making unconditional comparisons across different tool environments or reasoning tiers.
Applicable model versions: Claude Haiku 5.5; the page also lists Haiku 4.5, GPT-6 Luna, and Claude Sonnet 5.5 as comparisons.
Test environment or client: The Anthropic launch page lists GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 (Main), and Chartography. The page does not disclose the complete runtime environment.
Reasoning tier and parameters: Haiku 5.5 supports adjustable effort, and the charts show Low, Med, High, Xhigh, and Max. The complete parameters corresponding to each result are not disclosed.
This is an official benchmark summary on Anthropic's model launch page, not an independent rerun. The page lists model scores for knowledge work, computer use, multidisciplinary reasoning, agentic coding, and visual reasoning, and specifies conditions for some tests:
OSWorld 2.1 results are marked Offline subset. The page says the benchmark measures an agent's ability to complete long, multi-step tasks on a real computer.
Humanity's Last Exam reports both no tools and with tools results.
FrontierCode 1.1 uses the page's labeled Main subset; the comparison table marks the Sonnet 5.5 result as Xhigh.
The page also provides score and per-run cost curves for OSWorld 2.1, GDPval-AA, and Humanity's Last Exam across different effort tiers. OSWorld's vertical axis is a partial-credit score, not a uniform accuracy metric. The page does not list every chart value point by point.
The page directs readers to the Haiku 5.5 System Card for full evaluation details. This note collects only what the launch page shows directly.
| Evaluation category | Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|---|
| Knowledge work | GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| Knowledge work | AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| Computer use | OSWorld 2.1 (Offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Multidisciplinary reasoning | Humanity's Last Exam (no tools) | 45.9% | 10.2% | — | 56.9% |
| Multidisciplinary reasoning | Humanity's Last Exam (with tools) | 57.4% | 18.7% | — | 64.5% |
| Agentic coding | Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| Agentic coding | FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1% (Xhigh) |
| Visual reasoning | Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
The page also gives cost and capability curves, but does not provide a complete table of values for every effort tier in the body text. Its description of OSWorld 2.1 is that the evaluation measures an agent's ability to complete long, multi-step tasks on a real computer.
The Anthropic launch page gives the following prices per million tokens. Haiku 5.5 input, output, and cache-read prices are split into prompts at or below and above 100k tokens:
| Item | Haiku 5.5 (≤100k / >100k) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
| Input tokens | $0.10 / $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 / $2.50 | $5.00 | $10.00 |
The page says Haiku 5.5 is priced 90% below Haiku 4.5 for requests at or below 100k tokens and 50% below it for requests above 100k. Its footnote says about 90% of Haiku 4.5 requests fall in the lower tier; the newer tokenizer also changes token usage per task. Anthropic therefore estimates about 75% lower average running cost. It calls Haiku 5.5 its fastest model at standard model speeds, while Opus Fast Mode may be faster, and recommends Haiku for narrower work such as subagents, compression, and summarization when a larger model handles the primary task.
Benchmark names: GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 (Main), and Chartography.
Haiku 5.5 values: 1620, 1578, 72.4%, 45.9% (no tools), 57.4% (with tools), 39.2%, 46.4%, and 46.4%.
Conditions explicitly stated on the page: OSWorld 2.1 is the Offline subset; Humanity's Last Exam is split into no tools / with tools; Chartography is no tools; the Sonnet 5.5 FrontierCode result is marked Xhigh.
Cost data: Haiku 5.5 cache reads are $0.01 / $0.05, cache writes are $0.125 / $0.625, input is $0.10 / $0.50, and output is $0.50 / $2.50. All prices are per million tokens, with the values before and after the slash corresponding to prompts at or below and above 100k tokens.
Undisclosed fields: The launch page does not provide the sample count for each benchmark, per-question inputs, random seeds, complete system prompts, tool implementations, hardware, token limits, run counts, confidence intervals, or raw logs.
The launch page supports the conclusion that Haiku 5.5 improved clearly over Haiku 4.5 on this set of official benchmarks listed by Anthropic, and that the company positions it at the intersection of speed, price, and high-frequency narrow tasks. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2%, below Sonnet 5.5 at 70.6% in the same table, so the page itself presents Sonnet 5.5 and Opus 5.5 as better choices for complex agentic coding.
These results are a snapshot from a vendor launch page. The page does not disclose enough information to reproduce the experiments independently, nor can it be used to estimate the target business's accuracy, latency, total cost, or stability. Different effort levels, tool conditions, data subsets, and undisclosed runtime settings can all affect the comparison; production selection still requires rerunning the evaluation with your own inputs, tools, and cost constraints.
Fix the model as Claude Haiku 5.5 and record the specific model ID, effort tier, tool configuration, and token budget used by the API or client.
Start with the OSWorld 2.1, GDPval-AA, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 (Main), and Chartography benchmarks listed on the page. Record the data subset and no tools / with tools conditions.
Save complete inputs, scoring scripts, run counts, failure types, input and output tokens, latency, and cost for each benchmark. Do not invent configurations that the page does not provide.
Compare Haiku 5.5 and the comparison models in the same harness, on the same data version, and at the same effort tier. Report results for different tool conditions separately.
Add your own summarization, classification, browser-operation, and subagent tasks for business acceptance. Do not treat scores from the official launch page as a production commitment.
Claude Haiku 5.5