Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaClaude Opus 4.7

Claude Opus 4.7: Vellum's Cross-model Benchmarks and Task Selection

Original source

Vellum

AuthorNicolas Zeeb

Source date2026-04-16

Tabbit curation2026-08-19

Read original

One-sentence takeaway

Vellum's synthesis of the official data shows that Opus 4.7's strengths are concentrated in SWE-bench Pro, MCP-Atlas, Finance Agent, and visual reasoning, while BrowseComp is a relative regression point. Model selection should therefore be based on the workflow rather than the overall leaderboard.

Use cases

  • Suitable tasks: Real-world repository fixes, multi-tool orchestration, financial analysis, desktop operation, and complex chart comprehension.

  • Unsuitable tasks: Applications focused mainly on multi-page web search and synthesis, or applications with strict token budgets that have not yet remeasured the 4.7 tokenizer.

  • Applicable model version: Claude Opus 4.7.

  • Applicable client, Agent, or API: Vellum's analysis and the Anthropic API/platform; the article is not a controlled experiment on any one customer's production traffic.

  • Recommended reasoning tier and parameters: Start with high/xhigh as Anthropic recommends; the article itself does not independently disclose a standardized effort configuration.

Test environment, inputs/configuration

  • Data sources: Vellum compiled the benchmark table from Anthropic's official system card and partner materials, and explained what each benchmark measures.

  • Compared models: Opus 4.7/4.6, Claude Mythos Preview, GPT-5.4/5.4 Pro, and Gemini 3.1 Pro.

  • Original inputs and harness: Vellum did not publish the complete inputs, code, or run configuration used to rerun these benchmarks; this is an analysis of the data and of task selection.

Results data

BenchmarkOpus 4.7Opus 4.6Key comparison
SWE-bench Verified87.6%80.8%Gemini 3.1 Pro 80.6%
SWE-bench Pro64.3%53.4%GPT-5.4 57.7%; Gemini 3.1 Pro 54.2%
Terminal-Bench 2.069.4%65.4%GPT-5.4 75.1%; Gemini 3.1 Pro 68.5%
MCP-Atlas77.3%75.8%GPT-5.4 68.1%; Gemini 3.1 Pro 73.9%
Finance Agent v1.164.4%60.1%GPT-5.4 Pro 61.5%; Gemini 3.1 Pro 59.7%
OSWorld-Verified78.0%72.7%GPT-5.4 75.0%
BrowseComp79.3%83.7%GPT-5.4 Pro 89.3%; Gemini 3.1 Pro 85.9%
GPQA Diamond94.2%91.3%Gemini 3.1 Pro 94.3%
CharXiv without tools/with tools82.1% / 91.0%69.1% / 84.7%Mythos Preview 86.1% / 93.2%

Vellum also records the official migration note: the 4.7 tokenizer may turn the same text into approximately 1.0–1.35× as many tokens; long Agent turns at high effort may also increase output tokens.

Conclusion

If a workflow's bottleneck is complex coding, tool calls, screen understanding, or specialized analysis, Opus 4.7's gains are worth an A/B migration test. If the bottleneck is deep web search, its BrowseComp score of 79.3% is below 4.6's 83.7%, so consider retaining 4.6 or comparing other models.

Limitations

  • The article is primarily a second-order interpretation of official results, not an independent benchmark rerun by Vellum using the same harness.

  • Official partner figures (CursorBench, 93-task coding, visual acuity, and so on) lack the complete task definitions and raw outputs, so they should not be given equal weight with the public data.

  • The selection conclusions depend on the specific tools, context, effort, and token budget; the article does not provide a standardized API parameter table.

Reproduction steps

  1. Choose business samples from each of the four task categories listed in the article: repository issues, multi-tool workflows, visual charts, and web research.

  2. Hold the prompt, tool schema, permissions, effort, and context limit constant, then run Opus 4.6 and 4.7 separately.

  3. Record task success, tests passed, tool errors, tokens, latency, factual errors, and the amount of human revision.

  4. Examine BrowseComp-style research tasks separately so that coding scores do not average away a regression in web research.

  5. Estimate the per-task cost after the tokenizer change using real traffic, then decide whether to upgrade or route tasks across models.

Original evidence and data

The Vellum article explains SWE-bench Pro, MCP-Atlas, OSWorld, BrowseComp, GPQA, HLE, and CharXiv item by item, and explicitly notes that BrowseComp fell from 83.7% to 79.3%; the remaining table data comes from Anthropic's reported system card results.

Source excerpt or observation (short quote for compliance only)

Vellum's overall assessment is “a focused, high-performance upgrade,” rather than a sweeping win across the board; this is consistent with the split between the gains in coding/tool use and the decline in BrowseComp shown in the table.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Opus 4.7

Use and compare models in Tabbit

Claude Opus 4.7

Related reviews

MediaAnthropic Newsroom2026-04-16

Claude Opus 4.7: Official Coding, Vision, and Agent Benchmarks

CommunityReddit r/ClaudeCode

Claude Opus 4.7: Post-Release Long-Session Experience with Reddit Claude Code

Claude Opus 4.7

Related prompts

MediaAnthropic Claude Platform Docs / Anthropic Newsroom2026-04-16

Claude Opus 4.7: Effort Levels and Migration Prompt Template

MediaAnthropic official documentation

Claude Opus 4.7: Anthropic's Official Prompt Library and Best Patterns

MediaZooClaw AI Help2026-04-07

Claude Opus 4.7: Three Practical Patterns - Caveman Prompts, CLAUDE.md Safety Guardrails, and Git Worktrees

MediaTenten Learning

Claude Opus 4.7: The Complete Guide to 50+ Community Tips