Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 5 · Media / benchmark · Editorial analysis

Vellum Benchmark Cross-Comparison: Claude Sonnet 5 Six Major Benchmark Scores, Tokenizer Changes, and Cost Analysis

Sonnet 5, Vellum synthesis dated 2026-06-30; six benchmark families, 80.4% terminal control and 1,618 knowledge-work points; tasks and repeats differ.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
Sonnet 5, Vellum synthesis dated 2026-06-30; six benchmark families, 80.4% terminal control and 1,618 knowledge-work points; tasks and repeats differ

Key data and applicable tasks

One-sentence conclusion

Vellum's in-depth breakdown of Anthropic's official and third-party data shows that Sonnet 5 surpasses even the flagship Opus 4.8 in terminal control ( 80.4% ) and knowledge work ( 1,618 points ), but unit-task token expansion is pronounced under high effort and the updated Tokenizer, making model selection dependent on actual task costs.

Test environment

  • Benchmarks covered: SWE-Bench Pro ( coding ), Terminal-Bench 2.1 ( terminal interaction ), Humanity's Last Exam / HLE ( extreme reasoning ), OSWorld-Verified ( computer use ), GDPval-AA v2 ( professional knowledge work ), BrowseComp ( agentic search ).

  • Comparison models: Claude Sonnet 5 vs Claude Sonnet 4.6 vs Claude Opus 4.8.

  • Release context: Analysis of data and System Card technical details officially released by Anthropic.

Input/configuration

  • Standardized evaluation protocols disclosed in the official System Card ( including both tool-calling and tool-free versions ).

  • Evaluation of cost-performance curves across different adaptive thinking effort levels ( low, medium, high, xhigh ).

Results data

Evaluation BenchmarkClaude Sonnet 4.6Claude Sonnet 5Claude Opus 4.8Result Characteristics and Interpretation
Terminal-Bench 2.167.0%80.4%74.6%Outperforms Opus 4.8 ( 5.8% higher ), a 13.4% jump over Sonnet 4.6
GDPval-AA v2 ( Knowledge Work )-1,6181,615Knowledge work and analytical output are essentially on par with or slightly higher than Opus 4.8
Humanity's Last Exam ( with tools )46.8% ( corrected baseline )57.4%57.9%Nearing Opus 4.8, up 10.6% over Sonnet 4.6
OSWorld-Verified ( Computer Use )78.5% ( corrected baseline )81.2%83.4%Closely trailing Opus 4.8, offering more cost-effective graphical interface automation
SWE-Bench Pro ( Software Engineering )Baseline+5.0%+11.0%Positioned between 4.6 and 4.8, narrowing the gap
Firefox 147 Exploit ( Security Vulnerability )0.0% ( lower partial control )0.0% ( 13.2% partial control )Higher exploit rateOfficial confirmation of no specialized training on cyber exploits; hazardous capabilities remain contained

Conclusion

Sonnet 5 completely disrupts the conventional view that "mid-tier models consistently lag behind flagships," becoming the more cost-effective top choice for terminal operations, tool calling, and routine knowledge work tasks. However, model selection should not rely purely on per-token list prices: the updated Tokenizer introduces a 1.0–1.35x token expansion, and thinking token consumption surges under xhigh effort, making low/medium effort the primary sweet spot for deployment.

Limitations

  • In the HLE and OSWorld-Verified evaluations, Anthropic adjusted the historical baseline for Sonnet 4.6; historical comparisons must maintain consistent measurement criteria.

  • The token inflation effect results in higher actual dollar costs for long-text inputs compared to the legacy Sonnet 4.6.

Reproduction steps

  1. Configure uniform timeout, maximum token, and tool permission settings on a unified evaluation harness.

  2. Run the Terminal-Bench task suite across claude-sonnet-5, claude-sonnet-4-6, and claude-opus-4-8 respectively.

  3. Record the actual input character count, corresponding token count, thinking token percentage, and execution success rate for each task.

Original evidence and data

  • Core benchmark scores: Terminal-Bench 2.1 reached 80.4%, and GDPval-AA v2 scored 1,618 points.

  • Tokenizer mechanism: Sonnet 5 uses an updated Tokenizer; the same text input is tokenized into 1.0 to 1.35 times more tokens in Sonnet 5.

  • Cost boundaries: At low and medium effort levels, Sonnet 5's per-task cost at equivalent accuracy is noticeably lower than Opus 4.8; at xhigh effort, due to the heavy burn of thinking tokens, per-task cost can approach or even exceed that of Opus 4.8.

Source excerpts or observations ( for compliance short quotes only )

  • Source analysis: “Sonnet 5 doesn't close the gap to Opus 4.8 on terminal work. It moves past it... the first benchmark where the mid-tier model beats the flagship on the same harness.”

  • Source cost summary: “Sonnet 5 uses an updated tokenizer that maps the same input to 1.0–1.35x more tokens... Best value at low/medium effort; at xhigh it can cost more than Opus 4.8 for similar quality.”

What this supports

  • supports joint reading of capability, tokenizer expansion and task cost

What this does not support

  • does not combine heterogeneous benchmarks into a fixed win rate or infer your bill

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Vellum official blog · Vellum Team · Original publication date 2026-06-30 · Site edit date 2026-09-20

Open original source

Claude Sonnet 5

Compare Claude Sonnet 5 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 5: What Changed and How to Get Access

A sourced guide to Claude Sonnet 5, its changes from Sonnet 4.6, current access routes, limits, cost boundary and practical fit.

Related reviews

Claude Sonnet 5 Official Release: Agent Capabilities, Pricing Tiers, and Safety Boundaries2026-06-30 official release/system card; BrowseComp and OSWorld-Verified, coding, and safety tasks; prompts, repeats, and harness undisclosed.Endor Labs Independent Benchmark: Functional Correctness and Security Fix Performance of Claude Sonnet 5 with Claude CodeSonnet 5 with Claude Code, Agent Security League real vulnerability-fix tasks, report 2026-07-02; FuncPass 83.2%, SecPass 19.6%; no local rerun.CodeRabbit Production Field Report: In-Depth Comparison of Claude Sonnet 5 in Code Generation and PR Review QualitySonnet 5, CodeRabbit production review harness, report 2026-06-30; PR-comment precision about 38%–40%, sampling and repeats incomplete.Reddit community: Task experience and cost controversy after the Claude Sonnet 5 launchUser environment: Claude Max 5x, Claude's in-product memory and project context; specific API parameters, task sets, and tool harnesses were not disclosed consistently.。.Claude Sonnet 5 Official Prompting Methods: Effort Levels, Tool Calls, and Code ReviewSonnet 5 prompting focuses on using `effort` to control reasoning and cost first, then explicitly defining the task scope, tool-trigger conditions, and code-review phases.Reddit Community: Claude Sonnet 5 Response Truncation and Thinking Token Configuration Troubleshooting GuideTroubleshoot and resolve blank or mid-sentence cut-off responses in Claude Sonnet 5 across the API and third-party desktop clients ( Chatbox, AnythingLLM, etc. ) caused by adaptive thinking being enabled by default and exhausting `max_tokens`.Cursor Official Docs: Claude Sonnet 5 Model Integration, Usage Pools, and Agent Tool ConfigurationCursor positions Claude Sonnet 5 as the primary mid-tier coding model to replace Sonnet 4.6, supporting a 1M context window, thinking mode, and a complete Agent tool suite, with no long-context multiplier fees charged for contexts exceeding 200k tokens.Reddit Community: Tiered Model Routing with Opus Planning and Sonnet 5 Batch ExecutionEstablish a tiered division-of-labor workflow for complex projects—"flagship model ( Opus/Fable ) top-level planning + Sonnet 5 low/medium effort batch parallel execution + flagship model verification and synthesis"—to prevent Sonnet 5 from spinning its wheels across multiple turns and consuming excessive tokens on high-difficulty, open-ended tasks.