Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
CommunityClaude Sonnet 4.6

Harvey Legal Agent Bench: Sonnet 4.6 Full-Pass Rate 4.2%

Original source

X

AuthorHarvey (@harvey), citing Trajectory (@trajectorylabs)

Source date2026-06-11

Tabbit curation2026-08-20

Read original

One-sentence takeaway

Harvey recorded Claude Sonnet 4.6 at a 4.2% full-pass rate on the legal agent benchmark LAB, below Opus 4.6 at 6.6%; on the same leaderboard, post-trained NVIDIA Nemotron 3 Ultra reached 5.8%, and claimed operating costs are 1/8 to 1/50 of Sonnet/Opus.

Test environment

  • Benchmark: Harvey Legal Agent Bench (LAB).

  • Metric: full-pass / full-pass rate.

  • Comparison: Nemotron 3 Ultra baseline 0% → post-training 5.8%; Sonnet 4.6 4.2%; Opus 4.6 6.6%.

  • Additional observation: held-out tasks were ~70% pass before training (because enough scoring dimensions were missed), ~95% after training. The 70%/95% refers to Nemotron before/after post-training, not Sonnet.

Inputs/configuration

The X post did not disclose LAB questions, scoring dimensions, Sonnet's prompt, toolset, effort, or whether legal retrieval plugins were enabled. The Trajectory post stated the entire post-training was completed in under 24 hours after Nemotron 3 Ultra's release.

Results data

ModelLAB full-pass rate
Nemotron 3 Ultra (no post-training)0%
Claude Sonnet 4.64.2%
Nemotron 3 Ultra (legal post-training)5.8%
Claude Opus 4.66.6%

Harvey also wrote: the post-trained open-weight model reached quality close to leading closed-source models, with operating costs at 1/8 to 1/50 of Sonnet 4.6 and Opus 4.6 per-token prices.

Conclusion

Legal agent "full pass" is very strict: Sonnet 4.6's 4.2% does not mean it cannot do legal assistance, but means that under Harvey's all-dimensions-pass standard, it is clearly weaker than Opus 4.6 and also weaker than a specially post-trained Nemotron. Sonnet 4.6 is suitable for legal drafting/retrieval assistance, not as an automatic pass agent in LAB terms. High-risk legal deliverables should still go through Opus or domain post-trained models, with lawyer review retained.

Limitations

  • Full-pass rate is not partial correctness rate; 4.2% is not the same metric as everyday "can write a usable memo."

  • Per-question results and Sonnet configuration not disclosed.

  • Cost 1/8–1/50 is the author's statement on Nemotron vs Claude unit prices, not LAB scores.

  • Cannot conflate LAB with IDP document extraction or SWE coding.

Reproduction steps

  1. If Harvey/Trajectory later publish a LAB subset, fix the same scoring dimensions and run claude-sonnet-4-6 vs claude-opus-4-6.

  2. Report full pass, partial pass, citation errors, and missed scoring dimensions separately.

  3. Record retrieval tools, jurisdictions, and whether internet access was allowed.

  4. Legal conclusions must be reviewed by qualified personnel; model scores cannot substitute.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Sonnet 4.6

Use and compare models in Tabbit

Claude Sonnet 4.6

Related reviews

MediaAnthropic News / Introducing Sonnet 4.62026-02-17

Claude Sonnet 4.6 Official Release: Coding, Computer Use, and Agent Benchmarks

MediaBenchLM2026-08-17

BenchLM's Public Evidence Ledger for Claude Sonnet 4.6

MediaIDP Leaderboard

IDP Leaderboard: Sonnet 4.6 Matches Opus 4.6 on Real-World Document Understanding

MediaArtificial Analysis

Artificial Analysis: Sonnet 4.6 Non-Reasoning Intelligence Index 37

Claude Sonnet 4.6

Related prompts

MediaClaude Platform Docs / Prompting best practices

Clear Instructions, XML Context, and Self-Checking for Claude Sonnet 4.6

MediaClaude Platform Docs / Effort and Prompting best practices

Claude Sonnet 4.6: Effort and Tool-Triggering Configuration

MediaAnthropic Platform Docs / Computer Use API Reference2026-02-17

Claude Sonnet 4.6: Computer Use Tool Definitions and Automated Closed-Loop Workflow

MediaAnthropic Platform Docs / Context Management & Compaction2026-02-17

Claude Sonnet 4.6: 1M Long Context and Context Compaction Architecture Configuration