Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.
Tasks: BU Benchmark, Browser Use's benchmark for evaluating browser/web agents.
Comparison models: Gemini 3.6 Flash 68%; GPT-5.6-sol 67%; Claude Sonnet 4.6 62%; Claude Opus 4.8 74%.
Publication context: Released alongside Google DeepMind's Gemini 3.6 Flash launch, emphasizing web agent quality at Flash pricing.
Cost: The post says Gemini 3.6 Flash costs far less than Opus 4.8; no per-task dollar figures for Sonnet 4.6.
The X post does not disclose the task list, browser environment, step limits, whether screenshots were used, whether Sonnet had computer-use tools enabled, or effort/thinking settings. Scores are only comparable within Browser Use's harness.
| Model | BU Benchmark |
|---|---|
| Claude Opus 4.8 | 74% |
| Gemini 3.6 Flash | 68% |
| GPT-5.6-sol | 67% |
| Claude Sonnet 4.6 | 62% |
The author calls Gemini 3.6 Flash the most cost-effective browser agent model they have tested, writing that it ranks second only to Opus 4.8.
When you need a web agent, Sonnet 4.6 is a usable mid-tier option, but in Browser Use's public comparison it is neither the most accurate nor the cheapest. If the goal is pure browser automation, put Gemini 3.6 Flash and GPT-5.6-sol through the same harness before deciding; if the goal is desktop+browser hybrid computer use, this BU score cannot substitute for OSWorld.
Only four percentage figures, no per-task logs.
The vendor emphasized Flash when publishing competitor comparisons; Sonnet 4.6 was a baseline for comparison rather than an optimized target.
Does not specify Sonnet 4.6's API parameters, whether computer-use beta was used, or screenshot resolution.
Cannot be merged into one leaderboard with Anthropic's official OSWorld-Verified 72.5%.
Fix the same browser agent scaffold (or Browser Use's public benchmark), and connect claude-sonnet-4-6, Gemini 3.6 Flash, and GPT-5.6-sol separately.
Record success rate, steps, tokens, dollar cost, and types of pages where agents get stuck (login walls, dynamic DOM, multi-tab forms).
Separately track security failures such as prompt injection / mistaken payment clicks.
Do not add or subtract BU% and OSWorld%.
Claude Sonnet 4.6