This official guide describes Sonnet 5.5-specific prompting patterns for effort, carrying work through, running without up-front thinking, structured JSON, progress updates, tool use, mid-turn user messages, and coding verification. You can use them to adjust a Sonnet 5.5 system prompt and API harness.
Suitable tasks: Agentic coding, multi-turn tool use, research and support that require current information, multi-step reasoning with JSON output, progress updates during long tasks, and code changes that need a real test or build check.
Not suitable for: Treating an effort level with the same name as equivalent to the Sonnet 5 setting; increasing effort without measuring quality, latency, and cost; or using between_tools for a multi-step reasoning task with no tools.
Applicable model version: Claude Sonnet 5.5. The page says existing Claude Sonnet 5 prompts should usually work without changes; Opus is the better choice for the hardest long-horizon work.
Applicable client, agent, or API: Claude API, Claude Platform, and custom harnesses that provide tool use, streaming, and per-message effort control. Some settings are beta capabilities and should be checked against the current API documentation.
Recommended reasoning levels and parameters: Start at high, the API default; start at medium for well-specified agentic coding and multistep tool use, then move to high for harder or longer tasks; start at medium or low for chat and latency-sensitive work. For coding agents, leave room in max_tokens for thinking and the reply; the page recommends the model maximum of 128,000 and streaming the response.
The English text in the code fences below is reproduced from the source and can be reused directly. The surrounding explanations summarize the official applicability boundaries.
When the model often completes only part of a task and then asks what to do next at low or medium effort, add the following two paragraphs to the system prompt. The second paragraph can also be used on its own to limit unrequested additions.
Keep working until everything the user asked for is done, and only stop to ask when you can't go on without the user or before a risky step.
When the work the user asked for is done and checked, stop and report. Don't add features, tests, files, docs or refactors that weren't asked for. If you think one would help, mention it at the end instead of doing it.The official guide says the first paragraph makes low- and medium-effort sessions more likely to carry the work through, which increases duration and cost. It does not replace your own rules for risky actions.
When the harness provides subagents and xhigh or max effort starts extra review, hardening, or repair work on its own, use this system prompt fragment:
When the work the user asked for is done and its checks pass, stop and report. Don't start extra rounds of review or hardening on your own, and don't launch reviewer sub-agents unless the user asked for a review. If you think a deeper review is worth doing, say so at the end.The official guide says that in its coding-task tests, this reduced reviewer-subagent launches and cut session cost by about a third, with no measured change in quality. It does not eliminate self-started review rounds entirely.
For an open-ended request such as “show me what you can do,” if the user only wants ideas first, add this to the system prompt:
When the user asks for ideas, options or a plan, give them that and stop. Don't start building or changing anything until they say to go ahead.If the request already clearly asks for execution, say so directly in the user request so this rule is not applied too broadly.
When a task requires totaling figures, applying rules, or ranking items and the output must be JSON, the official guide recommends adaptive thinking and adding this line at the end of the system prompt:
Think the problem through before you answer.The guide explains that structured outputs can keep the response body limited to JSON that matches the schema, so the model must work through the problem in thinking. At high effort, this line brings accuracy close to xhigh with a modest increase in output tokens. When between_tools is used without tools, the model does not think first and this line has no effect.
If the product provides a search tool and the answer involves rules, requirements, prices, or other information that may change, use:
Use the search tool to check specifics that may have changed since your training, such as what is allowed, required or charged, even when you feel confident. For researched work such as a report or a comparison, gather current sources rather than writing from your training knowledge.Also remove older rules such as “only use tools when strictly necessary” or “minimize tool calls” if they suppress tool use.
When several consecutive tool calls produce no user-facing text or progress update, the harness can append a turn-scoped system message:
The user hasn't heard from you in a while — say in a few words what you're doing, then continue.The official guide recommends waiting until several consecutive silent steps, such as five, before sending a reminder; if the turn stays quiet, stop after the second or third reminder. Append the reminder after the latest tool results and keep it in later requests so the prompt cache and preserved thinking remain intact.
When the model reports that a code change is complete without running a test or build, add this paragraph to the system prompt:
When you change code that can be run, built, or type-checked, run a real check that exercises the change before reporting it done: the project's tests, type-checker, or build, or the changed command itself. A syntax-only check, or a check command that failed to start, does not count; if all that is missing is the project's declared dependencies, install them with its own package manager and lockfile (e.g. npm install, pip install -r requirements.txt), never via sudo or the system package manager, unless told not to. Only if no real check can run here, say which one you did not run and why instead of reporting the change as done.The official guide says that in its low-effort tests, this made skipped or superficial checks uncommon, with no measured change in task quality and only a small increase in per-task cost.
Sonnet 5.5 effort levels cannot be migrated by treating levels with the same names as equivalent to Sonnet 5. The official guide recommends a fresh sweep against your own evals. The page gives this configuration boundary:
{"type": "between_tools"}between_tools is Sonnet 5.5's lowest thinking setting and is accepted at high effort or below.
between_tools with xhigh or max returns a 400 error.
With between_tools, effort cannot be changed per message; use adaptive thinking for per-message effort changes.
For a request without tools that needs multistep reasoning, use adaptive thinking instead of between_tools.
Read responses by block type. Do not assume the first content block is text; pass the model's thinking blocks back unchanged.
Changing the top-level effort invalidates the prompt cache. For a one-turn change, use the supported beta per-message effort change and keep adaptive thinking enabled.
With structured outputs, low and medium effort can occasionally keep thinking until max_tokens and fail to finish normally. The official guide recommends:
Treat stop_reason: "max_tokens" as a failure even when the response text looks like valid JSON.
Leave enough max_tokens for thinking and the JSON, but no more than you are willing to spend on one attempt.
If structured outputs are unavailable, parse the last JSON value in the response instead of treating everything from the first { to the last } as JSON.
Validate the expected fields and retry once if necessary.
The guide's core approach is to adjust configuration based on the observed failure mode:
| Observed behavior | Official recommendation |
|---|---|
| Latency, quality, or token use changes after migrating from Sonnet 5 | Re-test effort instead of assuming same-named levels are equivalent |
| The agent stops to ask questions before finishing | Raise effort first, or add the “Keep working…” fragment |
| The agent adds unrequested tests, docs, or refactors | Use the second “done and checked” paragraph |
| xhigh / max starts reviewer agents on its own | Add the instruction to stop extra review, or lower effort |
| Unreliable JSON multi-step reasoning without tools | Use adaptive thinking and add “Think the problem through…” |
| A long tool chain appears to go silent | Enable display: "updates" beta, or add reminders after consecutive silent steps |
| Research answers use stale information | Provide a search tool and add the active-search fragment |
| Mid-turn user messages are mistaken for injection | Put user text in a user turn carrying the tool results, never inside tool_result; put harness notices in a separate system message |
| Code is reported complete without verification | Add the real test, build, or type-check fragment |
| Tool names or parameter names differ slightly | Accept an unambiguous match, or return an is_error: true tool result with the correct name |
| Dense charts or technical drawings lose detail | Provide crop, zoom, or code-execution image tools; for charts, tools help more than simply raising effort |
This guide works as a checklist for adjusting a Sonnet 5.5 system prompt and harness. The key migration points are to re-measure effort; control agent initiative and scope with explicit rules; use adaptive thinking for multi-step JSON reasoning; pass tool results, mid-turn user messages, and progress blocks using the correct message types; and run a real check before reporting a coding task complete.
The guide's figures such as “accuracy close to,” “about a third lower cost,” and “no measured change in quality” come from Anthropic's internal tests. It does not publish the complete task sets, sample sizes, model builds, prompt comparisons, or per-run outputs. These figures are useful for choosing what to test in a local eval, but they are not guarantees for an application.
You can build comparative tests from the source fragments, effort boundaries, message placement rules, and response-block handling guidance. To reproduce the guide's quality, cost, and latency claims, you would also need the exact Sonnet 5.5 deployment, the same task set, tool definitions, structured-output schema, effort, max_tokens, streaming configuration, retry policy, and scoring method. The guide does not publish all of these materials.
Claude Sonnet 5.5