BenchLM's displayable source ledger records 22 benchmark rows for Sonnet 4.6: 79.6% on SWE-bench Verified, 72.1% on OSWorld-Verified, 59.1% on Terminal-Bench, and 89.9% on GPQA. It also shows meaningful differences in strengths across categories, making it suitable as an entry point for verification rather than as a single overall score.
Data date: 2026-08-17; the page says it has 22 benchmark rows with displayable sources.
Aggregation: The page applies separate weights to categories including Coding, Agentic, Knowledge, Math, and Multimodal; the overall score is 64.53/100, ranked #39/218, but this score is BenchLM's custom aggregation.
Evidence markers: The table distinguishes Provider exact, Benchmark exact, and Secondary exact, and retains source links.
Verifiable sources: Rows for SWE-bench, OSWorld, and Terminal link to the benchmark or Anthropic system card.
This page is a ledger of results from multiple sources, not a single harness executed from scratch by BenchLM; its entries combine provider reports and benchmark leaderboards. When using it, open the source for each row, then fix the same benchmark, model snapshot, and configuration before comparing results.
Representative rows recorded on the page:
| Category/benchmark | Sonnet 4.6 score | Page evidence |
|---|---|---|
| SWE-bench Verified | 79.6% | Provider exact, linked to Anthropic system card |
| SWE-Rebench | 60.7% | Benchmark exact, SWE-Rebench leaderboard |
| OSWorld-Verified | 72.1% | Benchmark exact, OSWorld leaderboard |
| Terminal-Bench 2.0 | 59.1% | Provider exact, Anthropic system card |
| Claw-Eval | 67.8% | Benchmark exact |
| Humanity’s Last Exam | 49% | Provider exact, Anthropic system card |
| GPQA | 89.9% | Provider exact, Anthropic system card |
| SuperGPQA | 95% | Secondary exact |
| CharXiv Reasoning | 77.4% | Provider exact, Anthropic system card |
The displayable evidence supports the view that Sonnet 4.6 offers strong value in software engineering, computer agents, and knowledge question answering, but its custom overall score should not obscure clear differences on tasks such as Terminal-Bench and Math. During verification, prioritize item-by-item comparisons using the original benchmarks.
BenchLM's 64.53/100 and #39/218 are results from custom weights/directory listings, not ground truth for general capability.
The entries include provider reports, third-party leaderboards, and secondary sources, so their evidence levels differ.
The page's current directory includes future models and a dynamic leaderboard, so the results will change; record the collection date and link status.
For some rows, the model version, effort, tools, and number of evaluation repetitions are incomplete, so “fully reproducible” cannot be claimed unconditionally.
Open each source from the page, prioritizing official benchmark leaderboards or the Anthropic system card.
Fix the Sonnet 4.6 snapshot, effort, tool scaffold, task split, and timeout.
Rerun at least four task categories, including SWE, OSWorld, Terminal, and GPQA, and save the traces, errors, and costs.
Report raw scores separately from custom aggregation; do not use the BenchLM overall score as a substitute for task-level conclusions.
Claude Sonnet 4.6