Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 5 · Media / benchmark · Editorial analysis

Claude Sonnet 5 Official Release: Agent Capabilities, Pricing Tiers, and Safety Boundaries

2026-06-30 official release/system card; BrowseComp and OSWorld-Verified, coding, and safety tasks; prompts, repeats, and harness undisclosed.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkEditorial analysisEdited 2026-09-20

Test conditions

Conditions
2026-06-30 official release/system card; BrowseComp and OSWorld-Verified, coding, and safety tasks; prompts, repeats, and harness undisclosed

Key data and applicable tasks

Test environment

  • Model: Claude Sonnet 5, compared with Sonnet 4.6 and Opus 4.8.

  • Evaluation: The official presentation shows BrowseComp (Agentic Search) and OSWorld-Verified (computer use), and reports additional capability and safety evaluations in the System Card.

  • Inference control: effort tiers; official charts compare cost and performance at different effort levels.

  • Pricing: $2 per million input tokens and $10 per million output tokens; Anthropic says this introductory price became the permanent price on August 10.

Input/configuration

  • API model name: claude-sonnet-5.

  • Availability: The default model on Claude Free/Pro; also available on Max, Team, and Enterprise, and through the Claude API.

  • The agent evaluations use tool calls, long-running tasks, and different effort tiers; complete figures should be taken from the official charts and System Card.

Results

  • The official conclusion is that Sonnet 5 is clearly better than Sonnet 4.6 at agentic reasoning, tool use, coding, and knowledge work, with some high-effort tasks approaching Opus 4.8.

  • In the official cost-performance charts for BrowseComp and OSWorld-Verified, Sonnet 5 offers a broader range of cost options than Opus 4.8; Anthropic particularly emphasizes the cost efficiency of medium effort.

  • The official safety evaluation says that Sonnet 5's rates of harmful behavior, hallucination, sycophancy, and prompt-injection hijacking are generally lower than Sonnet 4.6's.

  • Cybersecurity boundary: In the Firefox exploit evaluation, neither Sonnet 5 nor Sonnet 4.6 completed a working exploit (both scored 0%); Sonnet 5's partial success rate was slightly higher. Network-security safeguards were therefore enabled at launch.

Conclusion

The official positioning is a “more agentic Sonnet”: it is suited to multi-step coding, tool calls, and browser/terminal workflows, while effort tiers provide room to choose among cost and completion-rate tradeoffs. Anthropic also explicitly states that its dangerous cyber capabilities are below those of Opus-level models, but that this does not justify omitting safety safeguards.

Limitations

  • The official materials consist of vendor testing and partner feedback, not independently reproduced experiments; specific configurations and complete scores for BrowseComp, OSWorld, and other evaluations should be taken from the System Card.

  • “Approaching Opus 4.8” depends on the task, effort, and budget, and cannot be generalized to equivalent performance across all text or coding tasks.

  • Pricing and default behavior may still change with the API product; at the time of collection, the page's 2026-08-10 update was used as the reference.

Reproduction steps

  1. Use the same task set and call Sonnet 5 separately with medium, high, and xhigh effort.

  2. Fix the tool definitions, timeout, maximum total tokens, and retry rules; record the success rate, tool calls, input/output tokens, and wall-clock time.

  3. Rerun Sonnet 4.6 and Opus 4.8 in the same harness, without mixing different system prompts or different tool permissions.

  4. For coding and cybersecurity tasks, separately retain records of safeguards, authorization, and human review; do not infer that a system is risk-free from its failure to complete an exploit.

Original evidence and data

  • Official API pricing: $2/M for input and $10/M for output; the official page says this price changed from an introductory price to a permanent price on August 10.

  • Anthropic describes Sonnet 5 as better able to continue executing, self-checking, and completing deliverables in complex multi-step tasks.

  • The official safety evaluation reports both a lower overall rate of harmful behavior and the boundary that network-security safeguards still need to be enabled.

Source excerpts or observations (for compliance short quotes only)

  • Official positioning phrase: “Our most agentic Sonnet yet”.

  • The official release page treats effort as a cost/performance control, rather than as a single fixed capability level.

What this supports

  • supports official agent-search, computer-use, coding and safety positioning

What this does not support

  • does not establish independent success, production stability, or Tabbit availability

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Anthropic official blog · Anthropic · Original publication date 2026-06-30 · Site edit date 2026-09-20

Open original source

Claude Sonnet 5

Compare Claude Sonnet 5 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 5: What Changed and How to Get Access

A sourced guide to Claude Sonnet 5, its changes from Sonnet 4.6, current access routes, limits, cost boundary and practical fit.

Related reviews

Endor Labs Independent Benchmark: Functional Correctness and Security Fix Performance of Claude Sonnet 5 with Claude CodeSonnet 5 with Claude Code, Agent Security League real vulnerability-fix tasks, report 2026-07-02; FuncPass 83.2%, SecPass 19.6%; no local rerun.CodeRabbit Production Field Report: In-Depth Comparison of Claude Sonnet 5 in Code Generation and PR Review QualitySonnet 5, CodeRabbit production review harness, report 2026-06-30; PR-comment precision about 38%–40%, sampling and repeats incomplete.Vellum Benchmark Cross-Comparison: Claude Sonnet 5 Six Major Benchmark Scores, Tokenizer Changes, and Cost AnalysisSonnet 5, Vellum synthesis dated 2026-06-30; six benchmark families, 80.4% terminal control and 1,618 knowledge-work points; tasks and repeats differ.Reddit community: Task experience and cost controversy after the Claude Sonnet 5 launchUser environment: Claude Max 5x, Claude's in-product memory and project context; specific API parameters, task sets, and tool harnesses were not disclosed consistently.。.Claude Sonnet 5 Official Prompting Methods: Effort Levels, Tool Calls, and Code ReviewSonnet 5 prompting focuses on using `effort` to control reasoning and cost first, then explicitly defining the task scope, tool-trigger conditions, and code-review phases.Reddit Community: Claude Sonnet 5 Response Truncation and Thinking Token Configuration Troubleshooting GuideTroubleshoot and resolve blank or mid-sentence cut-off responses in Claude Sonnet 5 across the API and third-party desktop clients ( Chatbox, AnythingLLM, etc. ) caused by adaptive thinking being enabled by default and exhausting `max_tokens`.Cursor Official Docs: Claude Sonnet 5 Model Integration, Usage Pools, and Agent Tool ConfigurationCursor positions Claude Sonnet 5 as the primary mid-tier coding model to replace Sonnet 4.6, supporting a 1M context window, thinking mode, and a complete Agent tool suite, with no long-context multiplier fees charged for contexts exceeding 200k tokens.Reddit Community: Tiered Model Routing with Opus Planning and Sonnet 5 Batch ExecutionEstablish a tiered division-of-labor workflow for complex projects—"flagship model ( Opus/Fable ) top-level planning + Sonnet 5 low/medium effort batch parallel execution + flagship model verification and synthesis"—to prevent Sonnet 5 from spinning its wheels across multiple turns and consuming excessive tokens on high-difficulty, open-ended tasks.