Anthropic positions Haiku 4.5 as a low-cost, high-speed model for real-time assistants, customer service, pair programming, Claude Code sub-agents, and computer use. It emphasizes coding quality close to Sonnet 4, while Sonnet 4.5 remains the frontier model for complex coding.
Model/entry points: Claude Haiku 4.5; Claude API, Claude Code, the app, Amazon Bedrock, and Google Vertex AI.
Cost: $1 per million input tokens and $5 per million output tokens.
Benchmarks and cases: SWE-bench Verified, Augment agentic coding, instruction-following for slide text, computer use, and internal safety evaluations; harnesses varied by project.
Safety: Anthropic published an ASL-2 classification and a system card link, stating that the model showed a lower rate of problematic behavior in automated alignment evaluations.
The release page does not disclose the complete prompts, model parameters, sample counts, repetition counts, or error details for each benchmark. It says that Haiku 4.5 reached 90% of Sonnet 4.5's performance in the Augment agentic coding evaluation, and cites product cases in hallucination, speed, and cost.
Anthropic characterizes Haiku 4.5 as providing a coding level similar to Sonnet 4, at about one-third the cost and more than twice the speed.
Augment agentic coding evaluation: Haiku 4.5 reached 90% of Sonnet 4.5's performance (as cited by the publisher).
Anthropic says Haiku 4.5 outperformed Sonnet 4 on computer-use tasks and was more responsive in real-time assistants, customer service, pair programming, and multi-agent projects.
The release page cites an instruction-following case for slide text: Haiku 4.5 at 65% versus 44% for the advanced model; the complete task set and scoring rules were not disclosed.
For short responses, batch extraction, real-time interaction, and divisible subtasks, Haiku 4.5's cost and latency advantages offer clear product value. Difficult planning, cross-file architecture, and critical decisions still require Sonnet/Opus or independent review.
All figures come from Anthropic or partner/customer cases and are official or partner-reported, with no complete public harness.
“90% of Sonnet 4.5” is relative performance, not a 90% absolute success rate.
Computer use and web tasks carry prompt-injection, permission, and accidental-action risks; the release page's safety evaluations do not replace production isolation.
Pricing and performance on the release date are not guaranteed to remain stable; model aliases, cloud platforms, and rate limits may change.
Fix the claude-haiku-4-5 snapshot, API provider, max_tokens, tool permissions, and task set.
Select five task categories—short-form Q&A, batch extraction, code completion, computer forms, and sub-agents—and compare them against Sonnet 4.5.
Record accuracy/per-question resolved status, p50/p95 latency, input and output tokens, call count, cost, and safety incidents.
Add escalation/human review to critical tasks, and report Haiku-only results separately from the final results after routing.
Claude Haiku 4.5