Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Media / benchmark · Vendor report

Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent Benchmarks

Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkVendor reportEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-02-17.
Harness/task
Model: Claude Sonnet 4.6, API entry point `claude-sonnet-4-6`; the 1M-token context window is in beta.; Pricing: The release page states a starting price of $3/$15 per million input/output tokens, the same as Sonnet 4.5.
Sample/gaps
Limitations noted: The 1M context window is in beta, and long-context pricing and platform limits may differ.; Computer use faces risks such as prompt injection; the release page only describes improvements in safety evaluations, so production deployments still require isolation, permissions, and handling of web content as untrusted.

Key data and applicable tasks

One-sentence takeaway

Anthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.

Test environment

  • Model: Claude Sonnet 4.6, API entry point claude-sonnet-4-6; the 1M-token context window is in beta.

  • Pricing: The release page states a starting price of $3/$15 per million input/output tokens, the same as Sonnet 4.5.

  • Benchmarks/workflows: SWE-bench Verified, OSWorld-Verified, Vending-Bench Arena, OfficeQA, Finance Agent v1.1, GDPval-AA, and Claude Code user preferences; the harness differs across metrics.

  • Version boundary: Anthropic specifically notes that Sonnet 4.5 and earlier used the original OSWorld, while Sonnet 4.5 onward uses OSWorld-Verified, so results cannot be compared directly across versions.

Inputs/configuration

The release page does not disclose the complete prompts, sampling parameters, number of repetitions, or per-question trajectories for all benchmarks. Claude Code user preferences come from early testing; Vending-Bench Arena is a simulated long-horizon management task; OSWorld-Verified uses simulated computers, mouse and keyboard, and real software scenarios.

Results data

  • In Claude Code, users preferred Sonnet 4.6 over Sonnet 4.5 about 70% of the time; versus Opus 4.5, users preferred it about 59% of the time.

  • In the data cited on the release page, Sonnet 4.6 matched Opus 4.6 on OfficeQA; Anthropic says it first reinvested in capability in Vending-Bench Arena, then turned toward profitability, and finished ahead.

  • Anthropic's customer/partner feedback says that Box improved by 15 percentage points over Sonnet 4.5 on reasoning-intensive question answering; one insurance benchmark reached 94%, but these cases do not disclose a complete, standardized methodology.

  • The release page says Sonnet 4.6 is significantly better at computer use than the previous generation and approaches a human-level experience on multi-tab spreadsheets and multi-step forms; it also acknowledges that the model still trails skilled humans.

Conclusion

Sonnet 4.6 is suitable for routing large volumes of routine coding, browser/desktop automation, enterprise document, and financial analysis tasks away from Opus to a cheaper default model; for deep refactoring, multi-agent coordination, and complex agentic search, run a Sonnet/Opus routing evaluation.

Limitations

  • The results are self-reported by the vendor, and many depend on Anthropic's own scaffolds, early user preferences, or customer cases; complete raw data and variance are not disclosed.

  • Commonly cited figures such as 79.6% on SWE and 72.5% on OSWorld should be checked against the corresponding configurations in the system card or release page; do not treat all benchmark results as a single leaderboard.

  • The 1M context window is in beta, and long-context pricing and platform limits may differ.

  • Computer use faces risks such as prompt injection; the release page only describes improvements in safety evaluations, so production deployments still require isolation, permissions, and handling of web content as untrusted.

Reproduction steps

  1. Fix the claude-sonnet-4-6 snapshot, effort, thinking, tool permissions, and context-window configuration.

  2. Build a consistent task set covering real code issues, browser forms, enterprise PDFs/spreadsheets, and financial analysis, and save all tool trajectories.

  3. Run Sonnet 4.6 and Opus 4.6 separately, reporting accuracy, resolved status, tool calls, latency, tokens, cost, and prompt-injection failures.

  4. Treat the official figures as baselines rather than reproduction results, and explain the differences between your harness and Anthropic's harness.

What this supports

  • Sonnet 4.6 is suitable for routing large volumes of routine coding, browser/desktop automation, enterprise document, and financial analysis tasks away from Opus to a cheaper default model; for deep refactoring, multi-agent coordination, and complex agentic search, run a Sonnet/Opus routing evaluation.

What this does not support

  • The 1M context window is in beta, and long-context pricing and platform limits may differ.
  • Computer use faces risks such as prompt injection; the release page only describes improvements in safety evaluations, so production deployments still require isolation, permissions, and handling of web content as untrusted.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Anthropic News / Introducing Sonnet 4.6 · Anthropic · Original publication date 2026-02-17 · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Browser Use BU Benchmark: Sonnet 4.6 Browser Agent 62%Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep AnalysisOn the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.Harvey Legal Agent Bench: Sonnet 4.6 Full-Pass Rate 4.2%Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.