Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Media / benchmark · Independent measurement

Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37

Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-08-20.
Harness/task
Display configuration: Claude Sonnet 4.6, Non-reasoning, Effort high.; Index version: Artificial Analysis Intelligence Index v4.1.1, including GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.
Sample/gaps
Limitations noted: Cost per task and verbosity are N/A; cannot infer per-task dollars from this page.; Intelligence Index is a composite of 9 items; cannot be decomposed into SWE or OSWorld.

Key data and applicable tasks

One-sentence takeaway

Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.

Test environment

  • Display configuration: Claude Sonnet 4.6, Non-reasoning, Effort high.

  • Index version: Artificial Analysis Intelligence Index v4.1.1, including GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.

  • Comparison pool: The page states non-reasoning models are compared only against other non-reasoning models; pricing tiers are compared against proprietary models in the same category with >$1/1M blended.

  • Maintenance status: This model is deprecated. Only continues default 10k input token load performance benchmarks; results for other loads are historical values. Suggests considering Claude Sonnet 5 (Non-reasoning).

Inputs/configuration

The page shows the non-reasoning + high effort variant, not the reasoning variant. Input modalities text+image, output text, context 1M. Specific per-benchmark prompts and harness are in AA's Intelligence Index methodology, not fully expanded on this model page.

Results data

Model summary at collection time:

MetricPage value
Intelligence Index37 (category #5 / 63; category median 23)
Speed46.1 output tok/s (category #36 / 63; page says notably slow)
Input price$3.00 / million tokens
Output price$15.00 / million tokens
Cache discount90%
Cost per Intelligence Index taskN/A
VerbosityN/A
Context1M tokens

Same-page Intelligence comparison bar (excerpt, higher is better): Claude Opus 5 (max) 63, Claude Fable 5 62, GPT-5.6 Sol (max) 61, …, Claude Sonnet 4.6 (Non-reasoning) 37. In the speed comparison, Sonnet 4.6 is 46 tok/s, slower than most listed comparison models.

Page copy: Among the leading non-reasoning models on intelligence, but relatively expensive for same-price-tier non-reasoning models; supports text/image input.

Conclusion

By AA's methodology, Sonnet 4.6 non-reasoning high effort remains clearly stronger than the category median, but is no longer at the intelligence frontier as of 2026-08, and speed is also on the slow side. Suitable as a production reference point for "known unit price, 1M window, non-reasoning," but not suitable as a proxy for the latest Sonnet anymore. When reasoning variant scores are needed, open AA's reasoning page; do not substitute this page's 37-point score.

Limitations

  • Page explicitly deprecated; some workloads no longer updated.

  • This page is Non-reasoning High Effort, not interchangeable with official default effort or results with thinking enabled.

  • Cost per task and verbosity are N/A; cannot infer per-task dollars from this page.

  • Intelligence Index is a composite of 9 items; cannot be decomposed into SWE or OSWorld.

Reproduction steps

  1. Fix AA's v4.1.1 methodology, explicitly choosing whether to run non-reasoning or reasoning.

  2. Set claude-sonnet-4-6 with effort=high, disable reasoning.

  3. Record Index subscores, tok/s, latency, and actual $3/$15 billing separately.

  4. When comparing against Sonnet 5 non-reasoning, use AA's currently maintained pages, not this deprecated page's historical loads.

What this supports

  • By AA's methodology, Sonnet 4.6 non-reasoning high effort remains clearly stronger than the category median, but is no longer at the intelligence frontier as of 2026-08, and speed is also on the slow side. Suitable as a production reference point for "known unit price, 1M window, non-reasoning," but not suitable as a proxy for the latest Sonnet anymore. When reasoning variant scores are needed, open AA's reasoning page; do not substitute this page's 37-point score.

What this does not support

  • Cost per task and verbosity are N/A; cannot infer per-task dollars from this page.
  • Intelligence Index is a composite of 9 items; cannot be decomposed into SWE or OSWorld.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Artificial Analysis · Artificial Analysis · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.BenchLM's Public Evidence Ledger for Claude Sonnet 4.6BenchLM's displayable source ledger records 22 benchmark rows for Sonnet 4.6: 79.6% on SWE-bench Verified, 72.1% on OSWorld-Verified, 59.1% on Terminal-Bench, and 89.9% on GPQA. It also shows meaningful differences in strengths across categories, making it suitable as an entry point for verification rather than as a single overall score.Reddit MLOps Observations on Task Tiering Between Claude Sonnet 4.6 and Opus 4.6The community attributes Sonnet 4.6's strengths to office work, finance, computer use, and routine coding, while viewing Opus 4.6 as stronger in deep reasoning, terminal coding, and agentic search. The post also explicitly warns that these are static benchmarks based on Anthropic's self-reported scaffolds.IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document UnderstandingOn the open document AI leaderboard, Claude Sonnet 4.6 scores 80.7 overall, slightly above Opus 4.6's 80.4, making Sonnet a good choice for offloading OCR, table extraction, layout understanding, and key information extraction from Opus; still watch for content moderation false positives on archived scans.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.