Claude Sonnet 4.6 · Media / benchmark · Vendor report
Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.
Model: Claude Sonnet 4.6, API entry point claude-sonnet-4-6; the 1M-token context window is in beta.
Pricing: The release page states a starting price of $3/$15 per million input/output tokens, the same as Sonnet 4.5.
Benchmarks/workflows: SWE-bench Verified, OSWorld-Verified, Vending-Bench Arena, OfficeQA, Finance Agent v1.1, GDPval-AA, and Claude Code user preferences; the harness differs across metrics.
Version boundary: Anthropic specifically notes that Sonnet 4.5 and earlier used the original OSWorld, while Sonnet 4.5 onward uses OSWorld-Verified, so results cannot be compared directly across versions.
The release page does not disclose the complete prompts, sampling parameters, number of repetitions, or per-question trajectories for all benchmarks. Claude Code user preferences come from early testing; Vending-Bench Arena is a simulated long-horizon management task; OSWorld-Verified uses simulated computers, mouse and keyboard, and real software scenarios.
In Claude Code, users preferred Sonnet 4.6 over Sonnet 4.5 about 70% of the time; versus Opus 4.5, users preferred it about 59% of the time.
In the data cited on the release page, Sonnet 4.6 matched Opus 4.6 on OfficeQA; Anthropic says it first reinvested in capability in Vending-Bench Arena, then turned toward profitability, and finished ahead.
Anthropic's customer/partner feedback says that Box improved by 15 percentage points over Sonnet 4.5 on reasoning-intensive question answering; one insurance benchmark reached 94%, but these cases do not disclose a complete, standardized methodology.
The release page says Sonnet 4.6 is significantly better at computer use than the previous generation and approaches a human-level experience on multi-tab spreadsheets and multi-step forms; it also acknowledges that the model still trails skilled humans.
Sonnet 4.6 is suitable for routing large volumes of routine coding, browser/desktop automation, enterprise document, and financial analysis tasks away from Opus to a cheaper default model; for deep refactoring, multi-agent coordination, and complex agentic search, run a Sonnet/Opus routing evaluation.
The results are self-reported by the vendor, and many depend on Anthropic's own scaffolds, early user preferences, or customer cases; complete raw data and variance are not disclosed.
Commonly cited figures such as 79.6% on SWE and 72.5% on OSWorld should be checked against the corresponding configurations in the system card or release page; do not treat all benchmark results as a single leaderboard.
The 1M context window is in beta, and long-context pricing and platform limits may differ.
Computer use faces risks such as prompt injection; the release page only describes improvements in safety evaluations, so production deployments still require isolation, permissions, and handling of web content as untrusted.
Fix the claude-sonnet-4-6 snapshot, effort, thinking, tool permissions, and context-window configuration.
Build a consistent task set covering real code issues, browser forms, enterprise PDFs/spreadsheets, and financial analysis, and save all tool trajectories.
Run Sonnet 4.6 and Opus 4.6 separately, reporting accuracy, resolved status, tool calls, latency, tokens, cost, and prompt-injection failures.
Treat the official figures as baselines rather than reproduction results, and explain the differences between your harness and Anthropic's harness.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Anthropic News / Introducing Sonnet 4.6 · Anthropic · Original publication date 2026-02-17 · Site edit date 2026-09-20
Open original sourceClaude Sonnet 4.6
Download the Tabbit client to check model access