Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityClaude Sonnet 4.6

OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep Analysis

Original source

OSWorld Benchmark Leaderboard / Evaluation Suite

AuthorOSWorld Evaluation Team & Independent Researchers

Source date2026-02-28

Tabbit curation2026-08-20

Read original

One-sentence takeaway

On the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.

Test environment

  • Model: Claude Sonnet 4.6 (API version claude-sonnet-4-6).

  • Pricing: Input $3.00 per million tokens, output $15.00 per million tokens.

  • Benchmark/workflow: OSWorld-Verified (369 cross-application real desktop tasks covering Chrome, LibreOffice Calc/Writer, VSCode, Thunderbird, GIMP, and native OS file management).

  • Version boundary: Evaluations uniformly run in real virtual desktop containers at 1024×768 resolution, without assistive accessibility label trees (pure pixel screenshot + coordinate click mode).

Inputs/configuration

  • Scaffold: Standard OSWorld agent loop.

  • Parameter configuration: effort=medium, max action steps limited to 25 steps/task, mandatory full-screen capture before each action with action trajectory recorded.

  • Task forms: e.g. "Aggregate sales data by condition in LibreOffice Calc and export as a CSV with a specific name", "Collect product information across multiple Chrome tabs and fill into a purchase form".

Results data

Task categorySonnet 4.6 success rateOpus 4.6 success rateSonnet 4.5 baselineHuman reference level
Full suite aggregate (OSWorld-Verified)72.5%72.7%61.4%72.36%
Web multi-step forms (Chrome)84.2%85.0%71.3%88.5%
Office documents and spreadsheets (LibreOffice)71.8%72.4%58.9%76.0%
File system and system configuration (OS/Terminal)79.5%78.9%69.2%82.0%
Complex professional software (GIMP/CAD)54.5%54.8%46.2%63.0%
Average tokens per task34.2k48.6k38.1k-
Average time per task (seconds)38.5s62.1s45.2s22.0s

Conclusion

  1. Cost-performance leap: Sonnet 4.6's overall computer use score differs from Opus 4.6 by less than 0.2%, but average completion time is 38% shorter and token cost is nearly 40% lower, making it highly practical for large-scale RPA and desktop agent deployment.

  2. Exceeding human norm: In standardized, well-structured cross-application data transcription tasks, Sonnet 4.6's success rate already slightly exceeds the normal human tester benchmark (72.5% vs 72.36%).

Limitations

  • Failure modes concentrated: Main failures concentrate on: 1) minor pixel-level dropdown arrow click offset (about 35% of failures); 2) action racing ahead due to slow asynchronous network loading; 3) software shortcut conflicts not triggered.

  • Advanced graphic editing weaker: In GIMP image cropping, layer blending, and other continuous spatial judgment tasks, success rate is only slightly above 50%.

Reproduction steps

  1. Clone the official evaluation repository https://github.com/xlang-ai/OSWorld.

  2. Configure environment variable ANTHROPIC_API_KEY, and in the config file specify model="claude-sonnet-4-6" and anthropic-beta="computer-use-2026-01-24".

  3. Run the evaluation command: python run.py --benchmark verified --model claude-sonnet-4-6 --max_steps 25.

  4. Export aggregated results and verify artifact file hashes and status using the evaluation script.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 4.6

Use and compare models in Tabbit

Claude Sonnet 4.6

Related reviews

MediaAnthropic News / Introducing Sonnet 4.62026-02-17

Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent Benchmarks

MediaBenchLM2026-08-17

BenchLM's Public Evidence Ledger for Claude Sonnet 4.6

MediaIDP Leaderboard

IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document Understanding

MediaArtificial Analysis

Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37

Claude Sonnet 4.6

Related prompts

MediaClaude Platform Docs / Prompting best practices

Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6

MediaClaude Platform Docs / Effort and Prompting best practices

Claude Sonnet 4.6: Effort and Tool-Triggering Configuration

MediaAnthropic Platform Docs / Computer Use API Reference2026-02-17

Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow

MediaAnthropic Platform Docs / Context Management & Compaction2026-02-17

Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration