Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Community source · Editorial analysis

Sansa Bench: Sonnet 4.6 High Reasoning Overall 0.733, Currently Ranked 4th

Sansa Bench's current overall leaderboard records Claude-Sonnet-4.6 Reasoning High at 0.733, tied with/just behind Gemini-3.1-Pro-Preview Reasoning Low, below Opus 5 / 4.8 high reasoning; it is suitable for assessing overall capability with reasoning enabled, not suitable for directly citing old scores from Reddit posts from six months ago.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceEditorial analysisEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-08-20.
Harness/task
Scale: Page states 103 models tested, Updated Aug 19, 2026.; Overall score: Equal-weight average across capabilities; scoring includes exact match, numeric match, code execution, LLM judges.
Sample/gaps
Limitations noted: The Reddit old post's 0.921 hallucination resistance and other breakdowns were not reproduced on the overall leaderboard's first screen on 2026-08-20; they cannot be treated as current official numbers.; LLM judges introduce evaluator model bias.

Key data and applicable tasks

One-sentence takeaway

Sansa Bench's current overall leaderboard records Claude-Sonnet-4.6 Reasoning High at 0.733, tied with/just behind Gemini-3.1-Pro-Preview Reasoning Low, below Opus 5 / 4.8 high reasoning; it is suitable for assessing overall capability with reasoning enabled, not suitable for directly citing old scores from Reddit posts from six months ago.

Test environment

  • Scale: Page states 103 models tested, Updated Aug 19, 2026.

  • Overall score: Equal-weight average across capabilities; scoring includes exact match, numeric match, code execution, LLM judges.

  • Current Top 5 (Overall):

    1. Claude-Opus-5 Reasoning High 0.798

    2. Claude-Opus-4.8 Reasoning High 0.785

    3. Gemini-3.1-Pro-Preview Reasoning High 0.747

    4. Claude-Sonnet-4.6 Reasoning High 0.733

    5. Gemini-3.1-Pro-Preview Reasoning Low 0.733

  • Method entry points: Inspiration & Acknowledgments, Compare Models, Model Analysis.

Inputs/configuration

The overall leaderboard displays the Reasoning High variant. The homepage does not list Sonnet 4.6 Reasoning None's current score on the first screen. Private question banks rotate, making it hard for models to game the benchmark; this also means Overall scores from different dates cannot be treated as the same test.

Results data

Current page can be cited directly:

  • Claude Sonnet 4.6 Reasoning High: Overall 0.733.

Historical snapshot (must not be mixed with the above): Around February 2026, an r/LLMDevs post claimed its private benchmark (later linked to the same Sansa methodology page) at that time showed:

  • Sonnet 4.6 reasoning off: 0.648, above GPT-5.2 low reasoning 0.604

  • Sonnet 4.6 high reasoning: 0.719, tied for top with Gemini 3 Pro Preview at the time, above GPT-5.2 high 0.649

  • hallucination resistance 0.921; social calibration 0.905; error detection 0.848

  • weaker in hard science than Gemini 3 Pro (philosophy 0.767 vs 0.900, chemistry 0.710 vs 0.839, economics 0.750 vs 0.812)

  • sycophancy resistance 0.716, below Sonnet 4.5 high's 0.755

Some in that Reddit thread questioned that the post "reads very much like AI" and reported many actual hallucinations; the author also attached https://trysansa.com/benchmark and https://docs.sansaml.com/. When citing historical breakdowns, you must label the post date, and the current overall leaderboard no longer shows 0.719.

Conclusion

With high reasoning enabled, Sonnet 4.6 still enters Sansa's overall top five, but newer Opus models have opened a gap. Suitable as a candidate for those who need reliability and can accept high reasoning cost. If reasoning is disabled, do not use the current homepage's 0.733; open that model's None/Low row separately or run the same methodology yourself.

Limitations

  • Question banks rotate; scores across months cannot be directly subtracted.

  • Homepage Top 5 only guarantees Reasoning High Overall; capability breakdowns should be taken from Model Analysis; this collection did not expand each capability table on the full page.

  • The Reddit old post's 0.921 hallucination resistance and other breakdowns were not reproduced on the overall leaderboard's first screen on 2026-08-20; they cannot be treated as current official numbers.

  • LLM judges introduce evaluator model bias.

Reproduction steps

  1. Open https://trysansa.com/benchmark, record the page Updated date and the model row's Reasoning tier.

  2. If breakdowns are needed, enter that model's Model Analysis and export hallucination, sycophancy, science, and other capabilities.

  3. Fix claude-sonnet-4-6 at reasoning off/high; do not mix with Opus 5's High on one "Sonnet capability" table.

  4. If comparing against the Reddit old post, separately create a "2026-02 snapshot" column; do not overwrite the current 0.733.

What this supports

  • With high reasoning enabled, Sonnet 4.6 still enters Sansa's overall top five, but newer Opus models have opened a gap. Suitable as a candidate for those who need reliability and can accept high reasoning cost. If reasoning is disabled, do not use the current homepage's 0.733; open that model's None/Low row separately or run the same methodology yourself.

What this does not support

  • The Reddit old post's 0.921 hallucination resistance and other breakdowns were not reproduced on the overall leaderboard's first screen on 2026-08-20; they cannot be treated as current official numbers.
  • LLM judges introduce evaluator model bias.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Sansa Bench (trysansa.com) · Sansa AI; Reddit historical snapshot from ExactMacaroon6673 · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.Reddit MLOps Observations on Task Tiering Between Claude Sonnet 4.6 and Opus 4.6The community attributes Sonnet 4.6's strengths to office work, finance, computer use, and routine coding, while viewing Opus 4.6 as stronger in deep reasoning, terminal coding, and agentic search. The post also explicitly warns that these are static benchmarks based on Anthropic's self-reported scaffolds.IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document UnderstandingOn the open document AI leaderboard, Claude Sonnet 4.6 scores 80.7 overall, slightly above Opus 4.6's 80.4, making Sonnet a good choice for offloading OCR, table extraction, layout understanding, and key information extraction from Opus; still watch for content moderation false positives on archived scans.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.