Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Community source · Editorial analysis

OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep Analysis

On the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceEditorial analysisEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-02-28.
Harness/task
Model: Claude Sonnet 4.6 (API version `claude-sonnet-4-6`).; Pricing: Input $3.00 per million tokens, output $15.00 per million tokens.
Sample/gaps
Limitations noted: Failure modes concentrated: Main failures concentrate on: 1) minor pixel-level dropdown arrow click offset (about 35% of failures); 2) action racing ahead due to slow asynchronous network loading; 3) software shortcut conflicts not triggered.; Advanced graphic editing weaker: In GIMP image cropping, layer blending, and other continuous spatial judgment tasks, success rate is only slightly above 50%.

Key data and applicable tasks

One-sentence takeaway

On the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.

Test environment

  • Model: Claude Sonnet 4.6 (API version claude-sonnet-4-6).

  • Pricing: Input $3.00 per million tokens, output $15.00 per million tokens.

  • Benchmark/workflow: OSWorld-Verified (369 cross-application real desktop tasks covering Chrome, LibreOffice Calc/Writer, VSCode, Thunderbird, GIMP, and native OS file management).

  • Version boundary: Evaluations uniformly run in real virtual desktop containers at 1024×768 resolution, without assistive accessibility label trees (pure pixel screenshot + coordinate click mode).

Inputs/configuration

  • Scaffold: Standard OSWorld agent loop.

  • Parameter configuration: effort=medium, max action steps limited to 25 steps/task, mandatory full-screen capture before each action with action trajectory recorded.

  • Task forms: e.g. "Aggregate sales data by condition in LibreOffice Calc and export as a CSV with a specific name", "Collect product information across multiple Chrome tabs and fill into a purchase form".

Results data

Task categorySonnet 4.6 success rateOpus 4.6 success rateSonnet 4.5 baselineHuman reference level
Full suite aggregate (OSWorld-Verified)72.5%72.7%61.4%72.36%
Web multi-step forms (Chrome)84.2%85.0%71.3%88.5%
Office documents and spreadsheets (LibreOffice)71.8%72.4%58.9%76.0%
File system and system configuration (OS/Terminal)79.5%78.9%69.2%82.0%
Complex professional software (GIMP/CAD)54.5%54.8%46.2%63.0%
Average tokens per task34.2k48.6k38.1k-
Average time per task (seconds)38.5s62.1s45.2s22.0s

Conclusion

  1. Cost-performance leap: Sonnet 4.6's overall computer use score differs from Opus 4.6 by less than 0.2%, but average completion time is 38% shorter and token cost is nearly 40% lower, making it highly practical for large-scale RPA and desktop agent deployment.

  2. Exceeding human norm: In standardized, well-structured cross-application data transcription tasks, Sonnet 4.6's success rate already slightly exceeds the normal human tester benchmark (72.5% vs 72.36%).

Limitations

  • Failure modes concentrated: Main failures concentrate on: 1) minor pixel-level dropdown arrow click offset (about 35% of failures); 2) action racing ahead due to slow asynchronous network loading; 3) software shortcut conflicts not triggered.

  • Advanced graphic editing weaker: In GIMP image cropping, layer blending, and other continuous spatial judgment tasks, success rate is only slightly above 50%.

Reproduction steps

  1. Clone the official evaluation repository https://github.com/xlang-ai/OSWorld.

  2. Configure environment variable ANTHROPIC_API_KEY, and in the config file specify model="claude-sonnet-4-6" and anthropic-beta="computer-use-2026-01-24".

  3. Run the evaluation command: python run.py --benchmark verified --model claude-sonnet-4-6 --max_steps 25.

  4. Export aggregated results and verify artifact file hashes and status using the evaluation script.

What this supports

  • Cost-performance leap: Sonnet 4.6's overall computer use score differs from Opus 4.6 by less than 0.2%, but average completion time is 38% shorter and token cost is nearly 40% lower, making it highly practical for large-scale RPA and desktop agent deployment.
  • Exceeding human norm: In standardized, well-structured cross-application data transcription tasks, Sonnet 4.6's success rate already slightly exceeds the normal human tester benchmark (72.5% vs 72.36%).

What this does not support

  • Failure modes concentrated: Main failures concentrate on: 1) minor pixel-level dropdown arrow click offset (about 35% of failures); 2) action racing ahead due to slow asynchronous network loading; 3) software shortcut conflicts not triggered.
  • Advanced graphic editing weaker: In GIMP image cropping, layer blending, and other continuous spatial judgment tasks, success rate is only slightly above 50%.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

OSWorld Benchmark Leaderboard / Evaluation Suite · OSWorld Evaluation Team & Independent Researchers · Original publication date 2026-02-28 · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent BenchmarksAnthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.Browser Use BU Benchmark: Sonnet 4.6 Browser Agent 62%Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.Harvey Legal Agent Bench: Sonnet 4.6 Full-Pass Rate 4.2%Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.