Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Community source · Editorial analysis

Browser Use BU Benchmark: Sonnet 4.6 Browser Agent 62%

Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceEditorial analysisEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-07-22.
Harness/task
Tasks: BU Benchmark, Browser Use's benchmark for evaluating browser/web agents.; Comparison models: Gemini 3.6 Flash 68%; GPT-5.6-sol 67%; Claude Sonnet 4.6 62%; Claude Opus 4.8 74%.
Sample/gaps
Limitations noted: Does not specify Sonnet 4.6's API parameters, whether computer-use beta was used, or screenshot resolution.; Cannot be merged into one leaderboard with Anthropic's official OSWorld-Verified 72.5%.

Key data and applicable tasks

One-sentence takeaway

Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.

Test environment

  • Tasks: BU Benchmark, Browser Use's benchmark for evaluating browser/web agents.

  • Comparison models: Gemini 3.6 Flash 68%; GPT-5.6-sol 67%; Claude Sonnet 4.6 62%; Claude Opus 4.8 74%.

  • Publication context: Released alongside Google DeepMind's Gemini 3.6 Flash launch, emphasizing web agent quality at Flash pricing.

  • Cost: The post says Gemini 3.6 Flash costs far less than Opus 4.8; no per-task dollar figures for Sonnet 4.6.

Inputs/configuration

The X post does not disclose the task list, browser environment, step limits, whether screenshots were used, whether Sonnet had computer-use tools enabled, or effort/thinking settings. Scores are only comparable within Browser Use's harness.

Results data

ModelBU Benchmark
Claude Opus 4.874%
Gemini 3.6 Flash68%
GPT-5.6-sol67%
Claude Sonnet 4.662%

The author calls Gemini 3.6 Flash the most cost-effective browser agent model they have tested, writing that it ranks second only to Opus 4.8.

Conclusion

When you need a web agent, Sonnet 4.6 is a usable mid-tier option, but in Browser Use's public comparison it is neither the most accurate nor the cheapest. If the goal is pure browser automation, put Gemini 3.6 Flash and GPT-5.6-sol through the same harness before deciding; if the goal is desktop+browser hybrid computer use, this BU score cannot substitute for OSWorld.

Limitations

  • Only four percentage figures, no per-task logs.

  • The vendor emphasized Flash when publishing competitor comparisons; Sonnet 4.6 was a baseline for comparison rather than an optimized target.

  • Does not specify Sonnet 4.6's API parameters, whether computer-use beta was used, or screenshot resolution.

  • Cannot be merged into one leaderboard with Anthropic's official OSWorld-Verified 72.5%.

Reproduction steps

  1. Fix the same browser agent scaffold (or Browser Use's public benchmark), and connect claude-sonnet-4-6, Gemini 3.6 Flash, and GPT-5.6-sol separately.

  2. Record success rate, steps, tokens, dollar cost, and types of pages where agents get stuck (login walls, dynamic DOM, multi-tab forms).

  3. Separately track security failures such as prompt injection / mistaken payment clicks.

  4. Do not add or subtract BU% and OSWorld%.

What this supports

  • When you need a web agent, Sonnet 4.6 is a usable mid-tier option, but in Browser Use's public comparison it is neither the most accurate nor the cheapest. If the goal is pure browser automation, put Gemini 3.6 Flash and GPT-5.6-sol through the same harness before deciding; if the goal is desktop+browser hybrid computer use, this BU score cannot substitute for OSWorld.

What this does not support

  • Does not specify Sonnet 4.6's API parameters, whether computer-use beta was used, or screenshot resolution.
  • Cannot be merged into one leaderboard with Anthropic's official OSWorld-Verified 72.5%.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · Browser Use (@browseruse) · Original publication date 2026-07-22 · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent BenchmarksAnthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep AnalysisOn the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.Harvey Legal Agent Bench: Sonnet 4.6 Full-Pass Rate 4.2%Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.