Vellum's in-depth breakdown of Anthropic's official and third-party data shows that Sonnet 5 surpasses even the flagship Opus 4.8 in terminal control ( 80.4% ) and knowledge work ( 1,618 points ), but unit-task token expansion is pronounced under high effort and the updated Tokenizer, making model selection dependent on actual task costs.
Benchmarks covered: SWE-Bench Pro ( coding ), Terminal-Bench 2.1 ( terminal interaction ), Humanity's Last Exam / HLE ( extreme reasoning ), OSWorld-Verified ( computer use ), GDPval-AA v2 ( professional knowledge work ), BrowseComp ( agentic search ).
Comparison models: Claude Sonnet 5 vs Claude Sonnet 4.6 vs Claude Opus 4.8.
Release context: Analysis of data and System Card technical details officially released by Anthropic.
Standardized evaluation protocols disclosed in the official System Card ( including both tool-calling and tool-free versions ).
Evaluation of cost-performance curves across different adaptive thinking effort levels ( low, medium, high, xhigh ).
| Evaluation Benchmark | Claude Sonnet 4.6 | Claude Sonnet 5 | Claude Opus 4.8 | Result Characteristics and Interpretation |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 67.0% | 80.4% | 74.6% | Outperforms Opus 4.8 ( 5.8% higher ), a 13.4% jump over Sonnet 4.6 |
| GDPval-AA v2 ( Knowledge Work ) | - | 1,618 | 1,615 | Knowledge work and analytical output are essentially on par with or slightly higher than Opus 4.8 |
| Humanity's Last Exam ( with tools ) | 46.8% ( corrected baseline ) | 57.4% | 57.9% | Nearing Opus 4.8, up 10.6% over Sonnet 4.6 |
| OSWorld-Verified ( Computer Use ) | 78.5% ( corrected baseline ) | 81.2% | 83.4% | Closely trailing Opus 4.8, offering more cost-effective graphical interface automation |
| SWE-Bench Pro ( Software Engineering ) | Baseline | +5.0% | +11.0% | Positioned between 4.6 and 4.8, narrowing the gap |
| Firefox 147 Exploit ( Security Vulnerability ) | 0.0% ( lower partial control ) | 0.0% ( 13.2% partial control ) | Higher exploit rate | Official confirmation of no specialized training on cyber exploits; hazardous capabilities remain contained |
Sonnet 5 completely disrupts the conventional view that "mid-tier models consistently lag behind flagships," becoming the more cost-effective top choice for terminal operations, tool calling, and routine knowledge work tasks. However, model selection should not rely purely on per-token list prices: the updated Tokenizer introduces a 1.0–1.35x token expansion, and thinking token consumption surges under xhigh effort, making low/medium effort the primary sweet spot for deployment.
In the HLE and OSWorld-Verified evaluations, Anthropic adjusted the historical baseline for Sonnet 4.6; historical comparisons must maintain consistent measurement criteria.
The token inflation effect results in higher actual dollar costs for long-text inputs compared to the legacy Sonnet 4.6.
Configure uniform timeout, maximum token, and tool permission settings on a unified evaluation harness.
Run the Terminal-Bench task suite across claude-sonnet-5, claude-sonnet-4-6, and claude-opus-4-8 respectively.
Record the actual input character count, corresponding token count, thinking token percentage, and execution success rate for each task.
Core benchmark scores: Terminal-Bench 2.1 reached 80.4%, and GDPval-AA v2 scored 1,618 points.
Tokenizer mechanism: Sonnet 5 uses an updated Tokenizer; the same text input is tokenized into 1.0 to 1.35 times more tokens in Sonnet 5.
Cost boundaries: At low and medium effort levels, Sonnet 5's per-task cost at equivalent accuracy is noticeably lower than Opus 4.8; at xhigh effort, due to the heavy burn of thinking tokens, per-task cost can approach or even exceed that of Opus 4.8.
Source analysis: “Sonnet 5 doesn't close the gap to Opus 4.8 on terminal work. It moves past it... the first benchmark where the mid-tier model beats the flagship on the same harness.”
Source cost summary: “Sonnet 5 uses an updated tokenizer that maps the same input to 1.0–1.35x more tokens... Best value at low/medium effort; at xhigh it can cost more than Opus 4.8 for similar quality.”
Claude Sonnet 5