Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Community source · Editorial analysis

IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document Understanding

On the open document AI leaderboard, Claude Sonnet 4.6 scores 80.7 overall, slightly above Opus 4.6's 80.4, making Sonnet a good choice for offloading OCR, table extraction, layout understanding, and key information extraction from Opus; still watch for content moderation false positives on archived scans.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceEditorial analysisEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-08-20.
Harness/task
Model: Leaderboard rank 8 Claude Sonnet 4.6, rank 9 Claude Opus 4.6, rank 16 Claude Haiku 4.5.; Task: Intelligent Document Processing, covering OCR, table extraction, key information extraction, and visual Q&A; the overall score is the mean of benchmark subscores.
Sample/gaps
Limitations noted: The homepage does not publish Claude's full sampling configuration; cost figures come from the Reddit sync post, not leaderboard main table fields.; Content moderation failures count toward relevant subscores; they do not mean the model "couldn't understand" the page.

Key data and applicable tasks

One-sentence takeaway

On the open document AI leaderboard, Claude Sonnet 4.6 scores 80.7 overall, slightly above Opus 4.6's 80.4, making Sonnet a good choice for offloading OCR, table extraction, layout understanding, and key information extraction from Opus; still watch for content moderation false positives on archived scans.

Test environment

  • Model: Leaderboard rank #8 Claude Sonnet 4.6, rank #9 Claude Opus 4.6, rank #16 Claude Haiku 4.5.

  • Task: Intelligent Document Processing, covering OCR, table extraction, key information extraction, and visual Q&A; the overall score is the mean of benchmark subscores.

  • Sub-benchmarks: OlmOCR, OmniDoc, IDP.

  • Scale: The Reddit sync post cites 16 models and 9000+ real documents; at collection time the main table had expanded to 26 models. Treat the leaderboard's current numbers as authoritative.

  • Artifacts: The site provides Code, Datasets, and Results Explorer; methodology links to GitHub.

Inputs/configuration

The homepage does not list the full prompt, temperature, effort, or whether thinking is enabled for each Claude model. For verification, open Results Explorer for per-document outputs and check the GitHub methodology for the harness.

Results data

Main table at collection time (Overall / OlmOCR / OmniDoc / IDP):

ModelOverallOlmOCROmniDocIDP
Nanonets OCR-385.987.490.080.2
GPT-5.483.581.085.384.4
Gemini-3-Pro82.877.788.881.8
Claude Sonnet 4.680.773.986.981.2
Claude Opus 4.680.474.185.981.1
Claude Haiku 4.571.261.279.672.9

The Reddit sync post at the time listed Sonnet 80.8 / Opus 80.3 / Haiku 69.6, and said extraction-task radar charts were nearly identical; Sonnet cost about $24/1K pages and Opus about $40/1K pages. Main table numbers have since been tweaked slightly; cite https://www.idp-leaderboard.org/ when quoting.

The sync post also noted that old newspaper scans, textbook pages, and historical documents sometimes trigger stricter Claude content moderation, mainly on OlmOCR and OmniDoc.

Conclusion

For document extraction, tables, and layout understanding, Sonnet 4.6 can replace Opus 4.6 as the default model; Haiku 4.5 is clearly behind. If the use case is specialized OCR, the current top entry is Nanonets OCR-3, not general-purpose Claude. Archived and historical scans require separate testing of moderation block rates.

Limitations

  • This is a document understanding leaderboard; do not extrapolate to SWE, computer use, or open-ended reasoning.

  • Reddit's older numbers differ from the main table at collection time by 0.1–1.6 points; cite the page snapshot at the time of reference.

  • The homepage does not publish Claude's full sampling configuration; cost figures come from the Reddit sync post, not leaderboard main table fields.

  • Content moderation failures count toward relevant subscores; they do not mean the model "couldn't understand" the page.

Reproduction steps

  1. Download the public datasets and evaluation code from the site, and fix the model ID claude-sonnet-4-6 and comparison models.

  2. Run OlmOCR, OmniDoc, and IDP per the GitHub methodology, and save per-page predictions.

  3. Report extraction accuracy, moderation block rate, and cost per thousand pages separately.

  4. Spot-check failed pages in Results Explorer to distinguish recognition errors from safety filtering.

What this supports

  • For document extraction, tables, and layout understanding, Sonnet 4.6 can replace Opus 4.6 as the default model; Haiku 4.5 is clearly behind. If the use case is specialized OCR, the current top entry is Nanonets OCR-3, not general-purpose Claude. Archived and historical scans require separate testing of moderation block rates.

What this does not support

  • The homepage does not publish Claude's full sampling configuration; cost figures come from the Reddit sync post, not leaderboard main table fields.
  • Content moderation failures count toward relevant subscores; they do not mean the model "couldn't understand" the page.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

IDP Leaderboard · IDP Leaderboard / shhdwi (Reddit sync note) · Original publication date Unknown · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37Artificial Analysis places Claude Sonnet 4.6 (Non-reasoning, High Effort) among comparable non-reasoning models at Intelligence Index 37, approximately 46 tok/s, input $3 / output $15 per million tokens, with a stated 1M context; the page also notes this model is deprecated, and the intelligence score no longer represents the latest Sonnet.Reddit MLOps Observations on Task Tiering Between Claude Sonnet 4.6 and Opus 4.6The community attributes Sonnet 4.6's strengths to office work, finance, computer use, and routine coding, while viewing Opus 4.6 as stronger in deep reasoning, terminal coding, and agentic search. The post also explicitly warns that these are static benchmarks based on Anthropic's self-reported scaffolds.CursorBench: Sonnet 4.6 Scores 49%, as a Baseline for Sonnet 5 Launch ComparisonCursor officially reported Claude Sonnet 5 at 57% and Claude Sonnet 4.6 at 49% on CursorBench; this shows 4.6 remains the comparison baseline for Cursor's internal coding agent evaluation, but the public post did not provide questions, configuration, or per-question trajectories.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.