Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Media / benchmark · Independent measurement

BenchLM's Public Evidence Ledger for Claude Sonnet 4.6

BenchLM's displayable source ledger records 22 benchmark rows for Sonnet 4.6: 79.6% on SWE-bench Verified, 72.1% on OSWorld-Verified, 59.1% on Terminal-Bench, and 89.9% on GPQA. It also shows meaningful differences in strengths across categories, making it suitable as an entry point for verification rather than as a single overall score.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-08-17.
Harness/task
Data date: 2026-08-17; the page says it has 22 benchmark rows with displayable sources.; Aggregation: The page applies separate weights to categories including Coding, Agentic, Knowledge, Math, and Multimodal; the overall score is 64.53/100, ranked 39/218, but this score is BenchLM's custom aggregation.
Sample/gaps
Limitations noted: The page's current directory includes future models and a dynamic leaderboard, so the results will change; record the collection date and link status.; For some rows, the model version, effort, tools, and number of evaluation repetitions are incomplete, so “fully reproducible” cannot be claimed unconditionally.

Key data and applicable tasks

One-sentence takeaway

BenchLM's displayable source ledger records 22 benchmark rows for Sonnet 4.6: 79.6% on SWE-bench Verified, 72.1% on OSWorld-Verified, 59.1% on Terminal-Bench, and 89.9% on GPQA. It also shows meaningful differences in strengths across categories, making it suitable as an entry point for verification rather than as a single overall score.

Test environment

  • Data date: 2026-08-17; the page says it has 22 benchmark rows with displayable sources.

  • Aggregation: The page applies separate weights to categories including Coding, Agentic, Knowledge, Math, and Multimodal; the overall score is 64.53/100, ranked #39/218, but this score is BenchLM's custom aggregation.

  • Evidence markers: The table distinguishes Provider exact, Benchmark exact, and Secondary exact, and retains source links.

  • Verifiable sources: Rows for SWE-bench, OSWorld, and Terminal link to the benchmark or Anthropic system card.

Inputs/configuration

This page is a ledger of results from multiple sources, not a single harness executed from scratch by BenchLM; its entries combine provider reports and benchmark leaderboards. When using it, open the source for each row, then fix the same benchmark, model snapshot, and configuration before comparing results.

Results data

Representative rows recorded on the page:

Category/benchmarkSonnet 4.6 scorePage evidence
SWE-bench Verified79.6%Provider exact, linked to Anthropic system card
SWE-Rebench60.7%Benchmark exact, SWE-Rebench leaderboard
OSWorld-Verified72.1%Benchmark exact, OSWorld leaderboard
Terminal-Bench 2.059.1%Provider exact, Anthropic system card
Claw-Eval67.8%Benchmark exact
Humanity’s Last Exam49%Provider exact, Anthropic system card
GPQA89.9%Provider exact, Anthropic system card
SuperGPQA95%Secondary exact
CharXiv Reasoning77.4%Provider exact, Anthropic system card

Conclusion

The displayable evidence supports the view that Sonnet 4.6 offers strong value in software engineering, computer agents, and knowledge question answering, but its custom overall score should not obscure clear differences on tasks such as Terminal-Bench and Math. During verification, prioritize item-by-item comparisons using the original benchmarks.

Limitations

  • BenchLM's 64.53/100 and #39/218 are results from custom weights/directory listings, not ground truth for general capability.

  • The entries include provider reports, third-party leaderboards, and secondary sources, so their evidence levels differ.

  • The page's current directory includes future models and a dynamic leaderboard, so the results will change; record the collection date and link status.

  • For some rows, the model version, effort, tools, and number of evaluation repetitions are incomplete, so “fully reproducible” cannot be claimed unconditionally.

Reproduction steps

  1. Open each source from the page, prioritizing official benchmark leaderboards or the Anthropic system card.

  2. Fix the Sonnet 4.6 snapshot, effort, tool scaffold, task split, and timeout.

  3. Rerun at least four task categories, including SWE, OSWorld, Terminal, and GPQA, and save the traces, errors, and costs.

  4. Report raw scores separately from custom aggregation; do not use the BenchLM overall score as a substitute for task-level conclusions.

What this supports

  • The displayable evidence supports the view that Sonnet 4.6 offers strong value in software engineering, computer agents, and knowledge question answering, but its custom overall score should not obscure clear differences on tasks such as Terminal-Bench and Math. During verification, prioritize item-by-item comparisons using the original benchmarks.

What this does not support

  • The page's current directory includes future models and a dynamic leaderboard, so the results will change; record the collection date and link status.
  • For some rows, the model version, effort, tools, and number of evaluation repetitions are incomplete, so “fully reproducible” cannot be claimed unconditionally.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM · BenchLM · Original publication date 2026-08-17 · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.Reddit MLOps Observations on Task Tiering Between Claude Sonnet 4.6 and Opus 4.6The community attributes Sonnet 4.6's strengths to office work, finance, computer use, and routine coding, while viewing Opus 4.6 as stronger in deep reasoning, terminal coding, and agentic search. The post also explicitly warns that these are static benchmarks based on Anthropic's self-reported scaffolds.IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document UnderstandingOn the open document AI leaderboard, Claude Sonnet 4.6 scores 80.7 overall, slightly above Opus 4.6's 80.4, making Sonnet a good choice for offloading OCR, table extraction, layout understanding, and key information extraction from Opus; still watch for content moderation false positives on archived scans.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.