Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityClaude Sonnet 4.6

Browser Use BU Benchmark: Sonnet 4.6 Browser Agent 62%

Original source

X

AuthorBrowser Use (@browseruse)

Source date2026-07-22

Tabbit curation2026-08-20

Read original

One-sentence takeaway

Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.

Test environment

  • Tasks: BU Benchmark, Browser Use's benchmark for evaluating browser/web agents.

  • Comparison models: Gemini 3.6 Flash 68%; GPT-5.6-sol 67%; Claude Sonnet 4.6 62%; Claude Opus 4.8 74%.

  • Publication context: Released alongside Google DeepMind's Gemini 3.6 Flash launch, emphasizing web agent quality at Flash pricing.

  • Cost: The post says Gemini 3.6 Flash costs far less than Opus 4.8; no per-task dollar figures for Sonnet 4.6.

Inputs/configuration

The X post does not disclose the task list, browser environment, step limits, whether screenshots were used, whether Sonnet had computer-use tools enabled, or effort/thinking settings. Scores are only comparable within Browser Use's harness.

Results data

ModelBU Benchmark
Claude Opus 4.874%
Gemini 3.6 Flash68%
GPT-5.6-sol67%
Claude Sonnet 4.662%

The author calls Gemini 3.6 Flash the most cost-effective browser agent model they have tested, writing that it ranks second only to Opus 4.8.

Conclusion

When you need a web agent, Sonnet 4.6 is a usable mid-tier option, but in Browser Use's public comparison it is neither the most accurate nor the cheapest. If the goal is pure browser automation, put Gemini 3.6 Flash and GPT-5.6-sol through the same harness before deciding; if the goal is desktop+browser hybrid computer use, this BU score cannot substitute for OSWorld.

Limitations

  • Only four percentage figures, no per-task logs.

  • The vendor emphasized Flash when publishing competitor comparisons; Sonnet 4.6 was a baseline for comparison rather than an optimized target.

  • Does not specify Sonnet 4.6's API parameters, whether computer-use beta was used, or screenshot resolution.

  • Cannot be merged into one leaderboard with Anthropic's official OSWorld-Verified 72.5%.

Reproduction steps

  1. Fix the same browser agent scaffold (or Browser Use's public benchmark), and connect claude-sonnet-4-6, Gemini 3.6 Flash, and GPT-5.6-sol separately.

  2. Record success rate, steps, tokens, dollar cost, and types of pages where agents get stuck (login walls, dynamic DOM, multi-tab forms).

  3. Separately track security failures such as prompt injection / mistaken payment clicks.

  4. Do not add or subtract BU% and OSWorld%.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 4.6

Use and compare models in Tabbit

Claude Sonnet 4.6

Related reviews

MediaAnthropic News / Introducing Sonnet 4.62026-02-17

Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent Benchmarks

MediaBenchLM2026-08-17

BenchLM's Public Evidence Ledger for Claude Sonnet 4.6

MediaIDP Leaderboard

IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document Understanding

MediaArtificial Analysis

Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37

Claude Sonnet 4.6

Related prompts

MediaClaude Platform Docs / Prompting best practices

Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6

MediaClaude Platform Docs / Effort and Prompting best practices

Claude Sonnet 4.6: Effort and Tool-Triggering Configuration

MediaAnthropic Platform Docs / Computer Use API Reference2026-02-17

Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow

MediaAnthropic Platform Docs / Context Management & Compaction2026-02-17

Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration