Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Opus 4.7 · Media / benchmark · Independent measurement

Claude Opus 4.7: Vellum's Cross-model Benchmarks and Task Selection

Vellum's synthesis of the official data shows that Opus 4.7's strengths are concentrated in SWE-bench Pro, MCP-Atlas, Finance Agent, and visual reasoning, while BrowseComp is a relative regression point. Model selection should therefore be based on the workflow rather than the overall leaderboard.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Model/version
Claude-Opus-4.7; source date: 2026-04-16.
Harness/task
Data sources: Vellum compiled the benchmark table from Anthropic's official system card and partner materials, and explained what each benchmark measures.; Compared models: Opus 4.7/4.6, Claude Mythos Preview, GPT-5.4/5.4 Pro, and Gemini 3.1 Pro.
Sample/gaps
Limitations noted: Official partner figures (CursorBench, 93-task coding, visual acuity, and so on) lack the complete task definitions and raw outputs, so they should not be given equal weight with the public data.; The selection conclusions depend on the specific tools, context, effort, and token budget; the article does not provide a standardized API parameter table.

Key data and applicable tasks

One-sentence takeaway

Vellum's synthesis of the official data shows that Opus 4.7's strengths are concentrated in SWE-bench Pro, MCP-Atlas, Finance Agent, and visual reasoning, while BrowseComp is a relative regression point. Model selection should therefore be based on the workflow rather than the overall leaderboard.

Use cases

  • Suitable tasks: Real-world repository fixes, multi-tool orchestration, financial analysis, desktop operation, and complex chart comprehension.

  • Unsuitable tasks: Applications focused mainly on multi-page web search and synthesis, or applications with strict token budgets that have not yet remeasured the 4.7 tokenizer.

  • Applicable model version: Claude Opus 4.7.

  • Applicable client, Agent, or API: Vellum's analysis and the Anthropic API/platform; the article is not a controlled experiment on any one customer's production traffic.

  • Recommended reasoning tier and parameters: Start with high/xhigh as Anthropic recommends; the article itself does not independently disclose a standardized effort configuration.

Test environment, inputs/configuration

  • Data sources: Vellum compiled the benchmark table from Anthropic's official system card and partner materials, and explained what each benchmark measures.

  • Compared models: Opus 4.7/4.6, Claude Mythos Preview, GPT-5.4/5.4 Pro, and Gemini 3.1 Pro.

  • Original inputs and harness: Vellum did not publish the complete inputs, code, or run configuration used to rerun these benchmarks; this is an analysis of the data and of task selection.

Results data

BenchmarkOpus 4.7Opus 4.6Key comparison
SWE-bench Verified87.6%80.8%Gemini 3.1 Pro 80.6%
SWE-bench Pro64.3%53.4%GPT-5.4 57.7%; Gemini 3.1 Pro 54.2%
Terminal-Bench 2.069.4%65.4%GPT-5.4 75.1%; Gemini 3.1 Pro 68.5%
MCP-Atlas77.3%75.8%GPT-5.4 68.1%; Gemini 3.1 Pro 73.9%
Finance Agent v1.164.4%60.1%GPT-5.4 Pro 61.5%; Gemini 3.1 Pro 59.7%
OSWorld-Verified78.0%72.7%GPT-5.4 75.0%
BrowseComp79.3%83.7%GPT-5.4 Pro 89.3%; Gemini 3.1 Pro 85.9%
GPQA Diamond94.2%91.3%Gemini 3.1 Pro 94.3%
CharXiv without tools/with tools82.1% / 91.0%69.1% / 84.7%Mythos Preview 86.1% / 93.2%

Vellum also records the official migration note: the 4.7 tokenizer may turn the same text into approximately 1.0–1.35× as many tokens; long Agent turns at high effort may also increase output tokens.

Conclusion

If a workflow's bottleneck is complex coding, tool calls, screen understanding, or specialized analysis, Opus 4.7's gains are worth an A/B migration test. If the bottleneck is deep web search, its BrowseComp score of 79.3% is below 4.6's 83.7%, so consider retaining 4.6 or comparing other models.

Limitations

  • The article is primarily a second-order interpretation of official results, not an independent benchmark rerun by Vellum using the same harness.

  • Official partner figures (CursorBench, 93-task coding, visual acuity, and so on) lack the complete task definitions and raw outputs, so they should not be given equal weight with the public data.

  • The selection conclusions depend on the specific tools, context, effort, and token budget; the article does not provide a standardized API parameter table.

Reproduction steps

  1. Choose business samples from each of the four task categories listed in the article: repository issues, multi-tool workflows, visual charts, and web research.

  2. Hold the prompt, tool schema, permissions, effort, and context limit constant, then run Opus 4.6 and 4.7 separately.

  3. Record task success, tests passed, tool errors, tokens, latency, factual errors, and the amount of human revision.

  4. Examine BrowseComp-style research tasks separately so that coding scores do not average away a regression in web research.

  5. Estimate the per-task cost after the tokenizer change using real traffic, then decide whether to upgrade or route tasks across models.

Original evidence and data

The Vellum article explains SWE-bench Pro, MCP-Atlas, OSWorld, BrowseComp, GPQA, HLE, and CharXiv item by item, and explicitly notes that BrowseComp fell from 83.7% to 79.3%; the remaining table data comes from Anthropic's reported system card results.

Source excerpt or observation (short quote for compliance only)

Vellum's overall assessment is “a focused, high-performance upgrade,” rather than a sweeping win across the board; this is consistent with the split between the gains in coding/tool use and the decline in BrowseComp shown in the table.

What this supports

  • If a workflow's bottleneck is complex coding, tool calls, screen understanding, or specialized analysis, Opus 4.7's gains are worth an A/B migration test. If the bottleneck is deep web search, its BrowseComp score of 79.3% is below 4.6's 83.7%, so consider retaining 4.6 or comparing other models.

What this does not support

  • Official partner figures (CursorBench, 93-task coding, visual acuity, and so on) lack the complete task definitions and raw outputs, so they should not be given equal weight with the public data.
  • The selection conclusions depend on the specific tools, context, effort, and token budget; the article does not provide a standardized API parameter table.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Vellum · Nicolas Zeeb · Original publication date 2026-04-16 · Site edit date 2026-09-20

Open original source

Claude Opus 4.7

Compare Claude Opus 4.7 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Opus 4.7: What Changed, Where It Fits, and When to Migrate

A sourced Claude Opus 4.7 overview covering the 4.6 upgrade, benchmark split, API access, cost, lifecycle and migration checks.

Related reviews

Claude Opus 4.7: Official Coding, Vision, and Agent BenchmarksThe official release positions Opus 4.7 as an upgrade over 4.6 for difficult software engineering, long-horizon Agents, and high-resolution vision, but its BrowseComp regression and higher token usage show that it is not an unconditional replacement for every task.Claude Opus 4.7: Post-Release Long-Session Experience with Reddit Claude CodeCommunity feedback suggests that Opus 4.7's long-session quality and perceived context retention vary widely: some users report a clear speedup on debugging and website tasks, while others encounter overcomplication, forgetting, hallucinations, and token/quota pressure. It therefore must be validated on your own Claude Code sessions.Claude Opus 4.7: Effort Levels and Migration Prompt TemplateFollow a task-specific guide for “Claude Opus 4.7: Effort Levels and Migration Prompt Template”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Opus 4.7: Anthropic's Official Prompt Library and Best PatternsFollow a task-specific guide for “Claude Opus 4.7: Anthropic's Official Prompt Library and Best Patterns”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Opus 4.7: Anthropic's Official Guide to Steering Claude Code - Choosing Among CLAUDE.md, Skills, Hooks, and SubagentsFollow a task-specific guide for “Claude Opus 4.7: Anthropic's Official Guide to Steering Claude Code - Choosing Among CLAUDE.md, Skills, Hooks, and Subagents”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Opus 4.7: Claude Code Cookbook Commands, Roles, and Automation ConfigurationFollow a task-specific guide for “Claude Opus 4.7: Claude Code Cookbook Commands, Roles, and Automation Configuration”; prerequisites, steps, checks, fixes, and source boundaries are explicit.