Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.
Anthropic News / Introducing Sonnet 4.6 · Read evidenceClaude Sonnet 4.6 · Reviews and evidence
Which Claude Sonnet 4.6 conclusions hold up?
Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.
This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.
Editorial takeaways
Editorial takeaways
On the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.
OSWorld Benchmark Leaderboard / Evaluation Suite · Read evidenceArtificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.
Artificial Analysis · Read evidenceFull reviews and related reading
Selected evidence
Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent Benchmarks
Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-02-17.
- Harness/task
- Model: Claude Sonnet 4.6, API entry point `claude-sonnet-4-6`; the 1M-token context window is in beta.; Pricing: The release page states a starting price of $3/$15 per million input/output tokens, the same as Sonnet 4.5.
- Sample/gaps
- Limitations noted: The 1M context window is in beta, and long-context pricing and platform limits may differ.; Computer use faces risks such as prompt injection; the release page only describes improvements in safety evaluations, so production deployments still require isolation, permissions, and handling of web content as untrusted.
OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep Analysis
On the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-02-28.
- Harness/task
- Model: Claude Sonnet 4.6 (API version `claude-sonnet-4-6`).; Pricing: Input $3.00 per million tokens, output $15.00 per million tokens.
- Sample/gaps
- Limitations noted: Failure modes concentrated: Main failures concentrate on: 1) minor pixel-level dropdown arrow click offset (about 35% of failures); 2) action racing ahead due to slow asynchronous network loading; 3) software shortcut conflicts not triggered.; Advanced graphic editing weaker: In GIMP image cropping, layer blending, and other continuous spatial judgment tasks, success rate is only slightly above 50%.
Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37
Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-08-20.
- Harness/task
- Display configuration: Claude Sonnet 4.6, Non-reasoning, Effort high.; Index version: Artificial Analysis Intelligence Index v4.1.1, including GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.
- Sample/gaps
- Limitations noted: Cost per task and verbosity are N/A; cannot infer per-task dollars from this page.; Intelligence Index is a composite of 9 items; cannot be decomposed into SWE or OSWorld.
Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus Planning
The OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-06.
- Harness/task
- Product: Claude Code (post flair: Question about Claude Code).; Configuration: The OP explicitly uses only Sonnet, and never above medium effort.
- Sample/gaps
- Limitations noted: The OP self-identifies as not the heaviest user but also says they use Claude "deeper than most"; neither claim can be verified.; The BrowseComp chart reading comes from another post's comments, with no original chart or official table; do not upgrade it to an independent review conclusion.
All sources
All sources
Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent Benchmarks
Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-02-17.
- Harness/task
- Model: Claude Sonnet 4.6, API entry point `claude-sonnet-4-6`; the 1M-token context window is in beta.; Pricing: The release page states a starting price of $3/$15 per million input/output tokens, the same as Sonnet 4.5.
- Sample/gaps
- Limitations noted: The 1M context window is in beta, and long-context pricing and platform limits may differ.; Computer use faces risks such as prompt injection; the release page only describes improvements in safety evaluations, so production deployments still require isolation, permissions, and handling of web content as untrusted.
OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep Analysis
On the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-02-28.
- Harness/task
- Model: Claude Sonnet 4.6 (API version `claude-sonnet-4-6`).; Pricing: Input $3.00 per million tokens, output $15.00 per million tokens.
- Sample/gaps
- Limitations noted: Failure modes concentrated: Main failures concentrate on: 1) minor pixel-level dropdown arrow click offset (about 35% of failures); 2) action racing ahead due to slow asynchronous network loading; 3) software shortcut conflicts not triggered.; Advanced graphic editing weaker: In GIMP image cropping, layer blending, and other continuous spatial judgment tasks, success rate is only slightly above 50%.
Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37
Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-08-20.
- Harness/task
- Display configuration: Claude Sonnet 4.6, Non-reasoning, Effort high.; Index version: Artificial Analysis Intelligence Index v4.1.1, including GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.
- Sample/gaps
- Limitations noted: Cost per task and verbosity are N/A; cannot infer per-task dollars from this page.; Intelligence Index is a composite of 9 items; cannot be decomposed into SWE or OSWorld.
Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus Planning
The OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-06.
- Harness/task
- Product: Claude Code (post flair: Question about Claude Code).; Configuration: The OP explicitly uses only Sonnet, and never above medium effort.
- Sample/gaps
- Limitations noted: The OP self-identifies as not the heaviest user but also says they use Claude "deeper than most"; neither claim can be verified.; The BrowseComp chart reading comes from another post's comments, with no original chart or official table; do not upgrade it to an independent review conclusion.
BenchLM's Public Evidence Ledger for Claude Sonnet 4.6
BenchLM's displayable source ledger records 22 benchmark rows for Sonnet 4.6: 79.6% on SWE-bench Verified, 72.1% on OSWorld-Verified, 59.1% on Terminal-Bench, and 89.9% on GPQA. It also shows meaningful differences in strengths across categories, making it suitable as an entry point for verification rather than as a single overall score.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-08-17.
- Harness/task
- Data date: 2026-08-17; the page says it has 22 benchmark rows with displayable sources.; Aggregation: The page applies separate weights to categories including Coding, Agentic, Knowledge, Math, and Multimodal; the overall score is 64.53/100, ranked 39/218, but this score is BenchLM's custom aggregation.
- Sample/gaps
- Limitations noted: The page's current directory includes future models and a dynamic leaderboard, so the results will change; record the collection date and link status.; For some rows, the model version, effort, tools, and number of evaluation repetitions are incomplete, so “fully reproducible” cannot be claimed unconditionally.
Reddit MLOps Observations on Task Tiering Between Claude Sonnet 4.6 and Opus 4.6
The community attributes Sonnet 4.6's strengths to office work, finance, computer use, and routine coding, while viewing Opus 4.6 as stronger in deep reasoning, terminal coding, and agentic search. The post also explicitly warns that these are static benchmarks based on Anthropic's self-reported scaffolds.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-02-17.
- Harness/task
- Environment: A Reddit user compiled material from Anthropic announcements, VentureBeat, TechCrunch, OfficeChai, and other sources.; Input/configuration: The post lists multiple release figures for Sonnet 4.6 and Opus 4.6, but provides no independently run inputs, repeat counts, or complete tool traces.
- Sample/gaps
- Limitations noted: Some figures in the post may differ across pages or versions and must not be treated as a real-time leaderboard.; There is no consistent scaffold, sample set, repetition count, or cost statistics; absolute success rates cannot be inferred.
IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document Understanding
On the open document AI leaderboard, Claude Sonnet 4.6 scores 80.7 overall, slightly above Opus 4.6's 80.4, making Sonnet a good choice for offloading OCR, table extraction, layout understanding, and key information extraction from Opus; still watch for content moderation false positives on archived scans.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-08-20.
- Harness/task
- Model: Leaderboard rank 8 Claude Sonnet 4.6, rank 9 Claude Opus 4.6, rank 16 Claude Haiku 4.5.; Task: Intelligent Document Processing, covering OCR, table extraction, key information extraction, and visual Q&A; the overall score is the mean of benchmark subscores.
- Sample/gaps
- Limitations noted: The homepage does not publish Claude's full sampling configuration; cost figures come from the Reddit sync post, not leaderboard main table fields.; Content moderation failures count toward relevant subscores; they do not mean the model "couldn't understand" the page.
CursorBench: Sonnet 4.6 Scores 49%, as a Baseline for Sonnet 5 Launch Comparison
Cursor officially reported Claude Sonnet 5 at 57% and Claude Sonnet 4.6 at 49% on CursorBench; this shows 4.6 remains the comparison baseline for Cursor's internal coding agent evaluation, but the public post did not provide questions, configuration, or per-question trajectories.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-07-01.
- Harness/task
- Task: CursorBench (Cursor's proprietary coding/agent evaluation; the post does not elaborate on the definition).; Comparison models: Claude Sonnet 5 vs Claude Sonnet 4.6.
- Sample/gaps
- Limitations noted: Vendor-provided comparisons when launching new products carry baseline-selection bias.; Whether 4.6 and 5 used the same effort, context window, and tool set was not stated.
Browser Use BU Benchmark: Sonnet 4.6 Browser Agent 62%
Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-07-22.
- Harness/task
- Tasks: BU Benchmark, Browser Use's benchmark for evaluating browser/web agents.; Comparison models: Gemini 3.6 Flash 68%; GPT-5.6-sol 67%; Claude Sonnet 4.6 62%; Claude Opus 4.8 74%.
- Sample/gaps
- Limitations noted: Does not specify Sonnet 4.6's API parameters, whether computer-use beta was used, or screenshot resolution.; Cannot be merged into one leaderboard with Anthropic's official OSWorld-Verified 72.5%.
Harvey Legal Agent Bench: Sonnet 4.6 Full-Pass Rate 4.2%
Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-06-11.
- Harness/task
- Benchmark: Harvey Legal Agent Bench (LAB).; Metric: full-pass / full-pass rate.
- Sample/gaps
- Limitations noted: Cost 1/8–1/50 is the author's statement on Nemotron vs Claude unit prices, not LAB scores.; Cannot conflate LAB with IDP document extraction or SWE coding.
Sansa Bench: Sonnet 4.6 High Reasoning Overall 0.733, Currently Ranked 4th
Sansa Bench's current overall leaderboard records Claude-Sonnet-4.6 Reasoning High at 0.733, tied with/just behind Gemini-3.1-Pro-Preview Reasoning Low, below Opus 5 / 4.8 high reasoning; it is suitable for assessing overall capability with reasoning enabled, not suitable for directly citing old scores from Reddit posts from six months ago.
Unverified: the original source could not be rechecked.
- Model/version
- Claude-Sonnet-4.6; source date: 2026-08-20.
- Harness/task
- Scale: Page states 103 models tested, Updated Aug 19, 2026.; Overall score: Equal-weight average across capabilities; scoring includes exact match, numeric match, code execution, LLM judges.
- Sample/gaps
- Limitations noted: The Reddit old post's 0.921 hallucination resistance and other breakdowns were not reproduced on the overall leaderboard's first screen on 2026-08-20; they cannot be treated as current official numbers.; LLM judges introduce evaluator model bias.
Claude Sonnet 4.6
Compare Claude Sonnet 4.6 in Tabbit
Model access, features, and permissions depend on your current client account.