Anthropic's current official configuration for Claude Sonnet 5.5 includes a 1M-token context window, a standard maximum output of 128K tokens, adaptive thinking with high as the default effort, $2 input pricing, and $10 output pricing; Bedrock is the only one of the five listed platforms whose model ID has an anthropic. prefix.
Suitable tasks: Claude API workflows that need fast responses, long context, text or image input, agentic tool use, and adjustable effort.
Not suitable for: Treating the standard 128K output limit and the Message Batches API's 300K beta limit as interchangeable; treating the model overview's high default effort as the default for every client; or sending non-default temperature, top_p, or top_k values.
Applicable model version: Claude Sonnet 5.5, with status Active (latest), using claude-sonnet-5-5 and the corresponding ID for each platform.
Applicable client, agent, or API: Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. The Message Batches API's 300K output is a separate beta capability that requires a beta header.
Recommended reasoning levels and parameters: The overview lists high as the default effort and says between_tools is available at high effort or below. Check the relevant API documentation for other effort levels and request fields, then balance quality, latency, and cost against the workload.
| Platform | Official model ID |
|---|---|
| Claude API | claude-sonnet-5-5 |
| Amazon Bedrock | anthropic.claude-sonnet-5-5 |
| Google Cloud | claude-sonnet-5-5 |
| Microsoft Foundry | claude-sonnet-5-5 |
| Claude Platform on AWS | claude-sonnet-5-5 |
The page uses claude-sonnet-5-5 as the main model-page ID and lists the anthropic. prefix as the Amazon Bedrock platform difference. Use the ID for the target platform when calling the model.
| Item | Official value |
|---|---|
| Context window | 1M tokens |
| Max output | 128K tokens |
| Max output (Message Batches API, beta) | 300K tokens |
| Thinking | Adaptive |
| Default effort | high |
| Comparative latency | Fast |
| Input → output | Text and images → text |
| Reliable knowledge cutoff | 2026-06 |
| Training data cutoff | 2026-06 |
| Item | Price (per million tokens) |
|---|---|
| Input | $2 |
| Output | $10 |
| 5-minute cache write | $2.50 |
| 1-hour cache write | $4 |
| Cache read | $0.10 |
| Message Batches API | 50% discount on input and output |
The model page says that the Message Batches API's 300K output limit requires the output-300k-2026-03-24 beta header. Batch discounts and the 300K output capability are API features and should not be applied to ordinary Messages API requests.
Sonnet 5.5 enables adaptive thinking by default, and the overview lists high as the default effort. The following is an incomplete illustrative request assembled from the model specifications, not an example from the overview. Check the current Messages API and effort documentation for exact request fields:
{
"model": "claude-sonnet-5-5",
"max_tokens": 128000,
"output_config": {
"effort": "high"
}
}max_tokens: 128000 represents the standard maximum output published on the page. A real request also needs the required Messages API fields such as messages. The page does not say that every request should use the maximum; set max_tokens according to the task and cost budget.
The model page says that Sonnet 5.5's lowest thinking setting is between_tools, available at high effort or below. It turns off up-front thinking while progress updates between tool calls still return as thinking blocks:
{
"thinking": {
"type": "between_tools"
},
"output_config": {
"effort": "medium"
}
}The overview does not provide a complete between_tools request example. It explicitly says the setting is available at high effort or below, so do not apply this illustration directly to higher effort levels. Check the current API documentation for other thinking request forms.
If you use the Message Batches API and a request needs more than the ordinary 128K output limit, the page lists this beta header:
output-300k-2026-03-24This records only the header name and the 300K capability listed on the page. Check the current API reference for the exact header-passing method, SDK version, and batch request structure.
Setting temperature, top_p, or top_k to a non-default value returns 400; do not carry old-model sampling parameters over directly.
The minimum cacheable prompt length is 512 tokens.
Adaptive thinking is enabled by default. To turn off up-front thinking, use between_tools and keep effort at high or below.
The page summary lists five migration points from Sonnet 5: between_tools, forced tool use, thinking blocks being tied to the model and conversation, the computer-use tool change on the Claude API and Google Cloud, and advisor-tool model compatibility.
One further change does not fail the request but changes the response shape: text between tool calls returns in thinking blocks. A client that renders only text blocks can appear silent during tool calls.
| Item | Official value |
|---|---|
| Status | Active (latest) |
| Release date | 2026-09-28 |
| Retirement | Not sooner than 2027-09-28 |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
The model page describes Sonnet 5.5 as a combination of speed and intelligence. Its comparison table lists a 1M-token context window, 128K maximum output, adaptive thinking, high default effort, and fast comparative latency. These are official model-overview labels, not an independent speed benchmark.
This overview works as a current configuration table for integrating Sonnet 5.5: choose the right model ID for the platform, then design requests around the 1M context window, 128K standard output, pricing, caching, and high default effort. Check the Message Batches API's 300K beta capability separately when longer batch output is needed.
The model page's Fast label, high default, adaptive thinking, pricing, and lifecycle are Anthropic's current official configuration claims. It does not provide a complete SDK example, platform-by-platform latency data, or runtime samples for every beta capability. Validate the max_tokens limit, between_tools, and non-default sampling-parameter rules with actual API requests and error handling.
Use the Models API or a platform console to verify model status, then send minimal requests to Claude API, Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS to confirm the model ID, default effort, context and output limits, cache pricing, and 400 errors for invalid parameters. To verify 300K batch output, also fix the Message Batches API, beta header, SDK version, and request size; the model overview alone does not establish these runtime behaviors.
Claude Sonnet 5.5