Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Sonnet 4.6 · Community source · Editorial analysis

Harvey Legal Agent Bench: Sonnet 4.6 Full-Pass Rate 4.2%

Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourceEditorial analysisEdited 2026-09-20

Test conditions

Model/version
Claude-Sonnet-4.6; source date: 2026-06-11.
Harness/task
Benchmark: Harvey Legal Agent Bench (LAB).; Metric: full-pass / full-pass rate.
Sample/gaps
Limitations noted: Cost 1/8–1/50 is the author's statement on Nemotron vs Claude unit prices, not LAB scores.; Cannot conflate LAB with IDP document extraction or SWE coding.

Key data and applicable tasks

One-sentence takeaway

Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.

Test environment

  • Benchmark: Harvey Legal Agent Bench (LAB).

  • Metric: full-pass / full-pass rate.

  • Comparison: Nemotron 3 Ultra baseline 0% → post-training 5.8%; Sonnet 4.6 4.2%; Opus 4.6 6.6%.

  • Additional observation: held-out tasks were ~70% pass before training (because enough scoring dimensions were missed), ~95% after training. The 70%/95% refers to Nemotron before/after post-training, not Sonnet.

Inputs/configuration

The X post did not disclose LAB questions, scoring dimensions, Sonnet's prompt, toolset, effort, or whether legal retrieval plugins were enabled. The Trajectory post stated the entire post-training was completed in under 24 hours after Nemotron 3 Ultra's release.

Results data

ModelLAB full-pass rate
Nemotron 3 Ultra (no post-training)0%
Claude Sonnet 4.64.2%
Nemotron 3 Ultra (legal post-training)5.8%
Claude Opus 4.66.6%

Harvey also wrote: the post-trained open-weight model reached quality close to leading closed-source models, with operating costs at 1/8 to 1/50 of Sonnet 4.6 and Opus 4.6 per-token prices.

Conclusion

Legal agent "full pass" is very strict: Sonnet 4.6's 4.2% does not mean it cannot do legal assistance, but means that under Harvey's all-dimensions-pass standard, it is clearly weaker than Opus 4.6 and also weaker than a specially post-trained Nemotron. Sonnet 4.6 is suitable for legal drafting/retrieval assistance, not as an automatic pass agent in LAB terms. High-risk legal deliverables should still go through Opus or domain post-trained models, with lawyer review retained.

Limitations

  • Full-pass rate is not partial correctness rate; 4.2% is not the same metric as everyday "can write a usable memo."

  • Per-question results and Sonnet configuration not disclosed.

  • Cost 1/8–1/50 is the author's statement on Nemotron vs Claude unit prices, not LAB scores.

  • Cannot conflate LAB with IDP document extraction or SWE coding.

Reproduction steps

  1. If Harvey/Trajectory later publish a LAB subset, fix the same scoring dimensions and run claude-sonnet-4-6 vs claude-opus-4-6.

  2. Report full pass, partial pass, citation errors, and missed scoring dimensions separately.

  3. Record retrieval tools, jurisdictions, and whether internet access was allowed.

  4. Legal conclusions must be reviewed by qualified personnel; model scores cannot substitute.

What this supports

  • Legal agent "full pass" is very strict: Sonnet 4.6's 4.2% does not mean it cannot do legal assistance, but means that under Harvey's all-dimensions-pass standard, it is clearly weaker than Opus 4.6 and also weaker than a specially post-trained Nemotron. Sonnet 4.6 is suitable for legal drafting/retrieval assistance, not as an automatic pass agent in LAB terms. High-risk legal deliverables should still go through Opus or domain post-trained models, with lawyer review retained.

What this does not support

  • Cost 1/8–1/50 is the author's statement on Nemotron vs Claude unit prices, not LAB scores.
  • Cannot conflate LAB with IDP document extraction or SWE coding.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · Harvey (@harvey), citing Trajectory (@trajectorylabs) · Original publication date 2026-06-11 · Site edit date 2026-09-20

Open original source

Claude Sonnet 4.6

Compare Claude Sonnet 4.6 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Sonnet 4.6: What It Is, Pricing, Access, and the Sonnet 5 Migration Question

A sourced overview of Claude Sonnet 4.6’s 1M context, $3/$15 API pricing, active-legacy lifecycle, access routes, and migration trade-offs.

Related reviews

OSWorld-Verified Independent Review: Claude Sonnet 4.6 Computer Use and GUI Task Deep AnalysisOn the standard Ubuntu desktop benchmark suite OSWorld-Verified, Claude Sonnet 4.6 achieved a 72.5% task success rate, nearly matching flagship Opus 4.6 (72.7%), but still shows some visual localization jitter in high-frequency complex dynamic popups and multi-level right-click menu scenarios.Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent BenchmarksAnthropic's release data positions Sonnet 4.6 as a lower-cost Opus-level candidate: it performs strongly on SWE-bench, OSWorld, OfficeQA, financial agents, and long contexts, but Opus 4.6 is still worth considering for extremely deep reasoning, complex search, and difficult refactoring.Browser Use BU Benchmark: Sonnet 4.6 Browser Agent 62%Browser Use scored Claude Sonnet 4.6 at 62% on its own BU Benchmark, below Gemini 3.6 Flash at 68%, GPT-5.6-sol at 67%, and Opus 4.8 at 74%; it shows 4.6 can work as a browser agent, but it is not the most cost-effective choice on that harness.Reddit: Sonnet 4.6 Medium Effort Handles Daily Work; Complex Projects Still Need Opus PlanningThe OP believes Sonnet 4.6 medium effort in Claude Code can already handle a large volume of daily and high-intensity tasks; the comment consensus is that simple execution can stay on Sonnet, while complex reasoning, planning, and high-pressure coding still require Opus for architecture first, then hand off to Sonnet for implementation.Claude Code: Sonnet 4.6 Engineering Architecture and Subagent DivisionFollow a task-specific guide for “Claude Code: Sonnet 4.6 Engineering Architecture and Subagent Division”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6Follow a task-specific guide for “Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop WorkflowFollow a task-specific guide for “Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture ConfigurationFollow a task-specific guide for “Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.