Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaClaude Sonnet 5

Vellum Benchmark Cross-Comparison: Claude Sonnet 5 Six Major Benchmark Scores, Tokenizer Changes, and Cost Analysis

Original source

Vellum official blog

AuthorVellum Team

Source date2026-06-30

Tabbit curation2026-08-20

Read original

One-sentence conclusion

Vellum's in-depth breakdown of Anthropic's official and third-party data shows that Sonnet 5 surpasses even the flagship Opus 4.8 in terminal control ( 80.4% ) and knowledge work ( 1,618 points ), but unit-task token expansion is pronounced under high effort and the updated Tokenizer, making model selection dependent on actual task costs.

Test environment

  • Benchmarks covered: SWE-Bench Pro ( coding ), Terminal-Bench 2.1 ( terminal interaction ), Humanity's Last Exam / HLE ( extreme reasoning ), OSWorld-Verified ( computer use ), GDPval-AA v2 ( professional knowledge work ), BrowseComp ( agentic search ).

  • Comparison models: Claude Sonnet 5 vs Claude Sonnet 4.6 vs Claude Opus 4.8.

  • Release context: Analysis of data and System Card technical details officially released by Anthropic.

Input/configuration

  • Standardized evaluation protocols disclosed in the official System Card ( including both tool-calling and tool-free versions ).

  • Evaluation of cost-performance curves across different adaptive thinking effort levels ( low, medium, high, xhigh ).

Results data

Evaluation BenchmarkClaude Sonnet 4.6Claude Sonnet 5Claude Opus 4.8Result Characteristics and Interpretation
Terminal-Bench 2.167.0%80.4%74.6%Outperforms Opus 4.8 ( 5.8% higher ), a 13.4% jump over Sonnet 4.6
GDPval-AA v2 ( Knowledge Work )-1,6181,615Knowledge work and analytical output are essentially on par with or slightly higher than Opus 4.8
Humanity's Last Exam ( with tools )46.8% ( corrected baseline )57.4%57.9%Nearing Opus 4.8, up 10.6% over Sonnet 4.6
OSWorld-Verified ( Computer Use )78.5% ( corrected baseline )81.2%83.4%Closely trailing Opus 4.8, offering more cost-effective graphical interface automation
SWE-Bench Pro ( Software Engineering )Baseline+5.0%+11.0%Positioned between 4.6 and 4.8, narrowing the gap
Firefox 147 Exploit ( Security Vulnerability )0.0% ( lower partial control )0.0% ( 13.2% partial control )Higher exploit rateOfficial confirmation of no specialized training on cyber exploits; hazardous capabilities remain contained

Conclusion

Sonnet 5 completely disrupts the conventional view that "mid-tier models consistently lag behind flagships," becoming the more cost-effective top choice for terminal operations, tool calling, and routine knowledge work tasks. However, model selection should not rely purely on per-token list prices: the updated Tokenizer introduces a 1.0–1.35x token expansion, and thinking token consumption surges under xhigh effort, making low/medium effort the primary sweet spot for deployment.

Limitations

  • In the HLE and OSWorld-Verified evaluations, Anthropic adjusted the historical baseline for Sonnet 4.6; historical comparisons must maintain consistent measurement criteria.

  • The token inflation effect results in higher actual dollar costs for long-text inputs compared to the legacy Sonnet 4.6.

Reproduction steps

  1. Configure uniform timeout, maximum token, and tool permission settings on a unified evaluation harness.

  2. Run the Terminal-Bench task suite across claude-sonnet-5, claude-sonnet-4-6, and claude-opus-4-8 respectively.

  3. Record the actual input character count, corresponding token count, thinking token percentage, and execution success rate for each task.

Original evidence and data

  • Core benchmark scores: Terminal-Bench 2.1 reached 80.4%, and GDPval-AA v2 scored 1,618 points.

  • Tokenizer mechanism: Sonnet 5 uses an updated Tokenizer; the same text input is tokenized into 1.0 to 1.35 times more tokens in Sonnet 5.

  • Cost boundaries: At low and medium effort levels, Sonnet 5's per-task cost at equivalent accuracy is noticeably lower than Opus 4.8; at xhigh effort, due to the heavy burn of thinking tokens, per-task cost can approach or even exceed that of Opus 4.8.

Source excerpts or observations ( for compliance short quotes only )

  • Source analysis: “Sonnet 5 doesn't close the gap to Opus 4.8 on terminal work. It moves past it... the first benchmark where the mid-tier model beats the flagship on the same harness.”

  • Source cost summary: “Sonnet 5 uses an updated tokenizer that maps the same input to 1.0–1.35x more tokens... Best value at low/medium effort; at xhigh it can cost more than Opus 4.8 for similar quality.”

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 5

Use and compare models in Tabbit

Claude Sonnet 5

Related reviews

MediaAnthropic official blog2026-06-30

Claude Sonnet 5 Official Release: Agent Capabilities, Pricing Tiers, and Safety Boundaries

CommunityReddit, r/ClaudeAI2026-06-30

Reddit community: Task experience and cost controversy after the Claude Sonnet 5 launch

MediaEndor Labs2026-07-02

Endor Labs Independent Benchmark: Functional Correctness and Security Fix Performance of Claude Sonnet 5 with Claude Code

MediaCodeRabbit official blog2026-06-30

CodeRabbit Production Field Report: In-Depth Comparison of Claude Sonnet 5 in Code Generation and PR Review Quality

Claude Sonnet 5

Related prompts

MediaAnthropic Claude Platform Docs2026-06-30

Claude Sonnet 5 Official Prompting Methods: Effort Levels, Tool Calls, and Code Review

MediaCursor Docs2026-07-01

Cursor Official Docs: Claude Sonnet 5 Model Integration, Usage Pools, and Agent Tool Configuration

CommunityReddit, r/claude2026-07-30

Reddit Community: Claude Sonnet 5 Response Truncation and Thinking Token Configuration Troubleshooting Guide

CommunityReddit, r/ClaudeAI2026-07-03

Reddit Community: Tiered Model Routing with Opus Planning and Sonnet 5 Batch Execution