The community attributes Sonnet 4.6's strengths to office work, finance, computer use, and routine coding, while viewing Opus 4.6 as stronger in deep reasoning, terminal coding, and agentic search. The post also explicitly warns that these are static benchmarks based on Anthropic's self-reported scaffolds.
Environment: A Reddit user compiled material from Anthropic announcements, VentureBeat, TechCrunch, OfficeChai, and other sources.
Input/configuration: The post lists multiple release figures for Sonnet 4.6 and Opus 4.6, but provides no independently run inputs, repeat counts, or complete tool traces.
Result format: A secondary compilation and opinion, not a controlled experiment.
The post does not disclose a complete, directly reusable prompt. Its reusable element is a routing hypothesis: try Sonnet first for production office work, finance, and computer use; route deep search, novel reasoning, and terminal coding to Opus; then validate the approach on the same task set.
The Anthropic figures cited in the post include: SWE-bench Verified, Sonnet 79.6% vs. Opus 80.8%; OSWorld-Verified, 72.5% vs. 72.7%; GDPval-AA Elo, 1633 vs. 1606; Finance Agent v1.1, 63.3% vs. 60.1%; GPQA, 89.9% vs. 91.3%; Terminal-Bench, 59.1% vs. 65.4%; BrowseComp, 74.7% vs. 84.0%; ARC-AGI-2, 58.3% vs. 68.8%.
The post also mentions pricing of $3/$15 for Sonnet 4.6 and $5/$25 for Opus 4.6, and notes that the 1M context window is in beta. These prices and versions should be checked against the current official page.
This community compilation can support a candidate strategy of routing by task and validating against evals: use Sonnet 4.6 as the default at scale, and route high-difficulty, long-horizon, or high-cost-of-error branches to Opus 4.6.
The data primarily relays official and media reporting; the author explicitly notes the lack of independent results from Aider, Chatbot Arena, and others.
Some figures in the post may differ across pages or versions and must not be treated as a real-time leaderboard.
There is no consistent scaffold, sample set, repetition count, or cost statistics; absolute success rates cannot be inferred.
Take real production tasks and stratify them into office work, finance, computer use, coding, search, and deep reasoning.
Hold the model version, tools, effort, timeout, and maximum token count constant, and run a blind test.
Record success rate, error severity, call count, latency, cost, and human preference; separately measure the gains from Sonnet→Opus routing.
For each new official or community benchmark, record its source, publication date, and harness; do not transfer static scores directly to production.
Claude Sonnet 4.6