The official model overview provides the current Claude Haiku 5.5 model IDs for each platform, a 1M-token context window, a 128K maximum output, input-length pricing tiers, and a default medium effort baseline for API calls.
Suitable tasks: Selecting a model for high-volume, low-latency classification, routing, information extraction, and subagent requests, and creating a unified configuration table for the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS.
Not suitable for: Using this page as a substitute for Haiku 4.5 migration compatibility checks, business-quality evaluation, or platform SDK error handling. See the official migration configuration guide for migration differences.
Applicable model versions: Claude Haiku 5.5. Most listed platforms use claude-haiku-5-5; Amazon Bedrock uses an ID with the anthropic. prefix.
Applicable clients, agents, or APIs: The Claude API Messages API and the Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS integrations listed on the page.
Recommended reasoning tier and parameters: Keep adaptive thinking enabled, with output_config.effort set to the default medium, then adjust effort based on quality, latency, and cost evaluation. The page also requires omitting temperature, top_p, and top_k, because non-default values return 400.
| Platform | Claude Haiku 5.5 model ID |
|---|---|
| Claude API | claude-haiku-5-5 |
| Amazon Bedrock | anthropic.claude-haiku-5-5 |
| Google Cloud | claude-haiku-5-5 |
| Microsoft Foundry | claude-haiku-5-5 |
| Claude Platform on AWS | claude-haiku-5-5 |
| Item | Official value |
|---|---|
| Context window | 1M tokens |
| Maximum output | 128K tokens |
| Message Batches API maximum output | 300K tokens (beta; requires the output-300k-2026-03-24 beta header) |
| thinking | Adaptive, enabled by default |
| Default effort | medium |
| Input and output types | Text and images → text |
| Comparative latency | Fastest |
| Knowledge cutoff | June 2026 |
| Training data cutoff | June 2026 |
The official page says to use the effort parameter to control adaptive thinking depth. max_tokens includes thinking tokens, so it cannot be sized only around visible text length; reserve a sufficient limit for the target task.
The page also says Haiku 5.5 uses a newer tokenizer, counting about 30% more tokens for the same text than Haiku 4.5. Recalculate length and cost when migrating. Thinking blocks can be used only in the account that generated them or a linked account; check that constraint before replaying a conversation across accounts.
All prices below are per million tokens (MTok):
| Item | Prompts up to 100,000 tokens | Prompts over 100,000 tokens |
|---|---|---|
| Input | $0.10 / MTok | $0.50 / MTok |
| Output | $0.50 / MTok | $2.50 / MTok |
| 5-minute cache write | $0.125 / MTok | $0.625 / MTok |
| 1-hour cache write | $0.20 / MTok | $1.00 / MTok |
| Cache read | $0.01 / MTok | $0.05 / MTok |
The Batch API offers a 50% discount on input and output. Verify the actual bill against the platform's price list and whether the request hits the cache.
The following is a minimal request skeleton based on the model ID, adaptive thinking, and default effort in the model overview. Replace the messages content with the business input; the page does not require a fixed max_tokens value, so this example shows only the configuration relationship.
import anthropic
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-haiku-5-5",
max_tokens=8192,
thinking={"type": "adaptive"},
output_config={"effort": "medium"},
messages=[
{"role": "user", "content": "Replace this with the business input"},
],
)Do not set temperature, top_p, or top_k in this request. If you use the 300K-output Batch API beta capability, add the output-300k-2026-03-24 beta header as described on the official page and confirm that the client supports this beta interface.
Choose the model ID from the table for the deployment platform, and record the platform, model ID, context limit, and pricing tiers in the configuration center.
Use adaptive thinking with effort=medium by default. Measure quality, time to first token, total latency, output tokens, and per-request cost separately for each business task.
Track the cost of requests over 100,000 input tokens separately; do not estimate long-context tasks with the lower tier.
If prompt caching is enabled, record 5-minute writes, 1-hour writes, and cache reads separately. Do not mistake the cache-read price for the ordinary input price.
If using the Batch API, confirm that the 50% input/output discount and the 300K-output beta header both apply to the current platform and SDK.
Query the Models API for the models, capabilities, and limits available to the account. Do not treat the status collected in this note as a permanent availability commitment.
This is the official model overview configuration collected by Anthropic on 2026-10-09. It is not an independent performance evaluation and does not guarantee a particular accuracy, latency, or cost for business tasks.
The page marks Claude Haiku 5.5 as Active (latest) and says it will not retire before 2027-10-07. Actual platform availability still depends on the account, region, quota, and provider release status.
“Fastest” is a relative latency description on the official page. It does not include standardized hardware, concurrency, input length, or statistical intervals, so it cannot be converted directly into a millisecond SLA.
Pricing is tiered based on whether the prompt exceeds 100,000 tokens. Thinking tokens, cache hits, the Batch API, and different cloud platforms' billing details should be checked against actual billing and the platform's price list.
This page records the current model page's call, capacity, and pricing baseline. Assistant prefill, tool sets, message replay, and other version-compatibility changes belong to migration work and should be checked item by item in the migration configuration guide.
Model IDs: The page's Specifications section lists claude-haiku-5-5 for the Claude API, anthropic.claude-haiku-5-5 for Amazon Bedrock, and claude-haiku-5-5 for the other listed platforms.
Capacity: The page's at-a-glance and Capabilities sections both list a 1M context window and 128K maximum output; the Message Batches API separately lists a 300K beta limit.
Pricing: The page lists input, output, 5-minute cache-write, 1-hour cache-write, and cache-read prices split at 100,000 tokens, and lists a 50% Batch API input/output discount.
effort: The page lists Thinking as Adaptive and Default effort as medium, and says that effort can control thinking depth.
Parameter limits: The page's Good to know section explicitly requires omitting temperature, top_p, and top_k, because setting non-default values returns 400.
Claude Haiku 5.5