Vellum's synthesis of the official data shows that Opus 4.7's strengths are concentrated in SWE-bench Pro, MCP-Atlas, Finance Agent, and visual reasoning, while BrowseComp is a relative regression point. Model selection should therefore be based on the workflow rather than the overall leaderboard.
Suitable tasks: Real-world repository fixes, multi-tool orchestration, financial analysis, desktop operation, and complex chart comprehension.
Unsuitable tasks: Applications focused mainly on multi-page web search and synthesis, or applications with strict token budgets that have not yet remeasured the 4.7 tokenizer.
Applicable model version: Claude Opus 4.7.
Applicable client, Agent, or API: Vellum's analysis and the Anthropic API/platform; the article is not a controlled experiment on any one customer's production traffic.
Recommended reasoning tier and parameters: Start with high/xhigh as Anthropic recommends; the article itself does not independently disclose a standardized effort configuration.
Data sources: Vellum compiled the benchmark table from Anthropic's official system card and partner materials, and explained what each benchmark measures.
Compared models: Opus 4.7/4.6, Claude Mythos Preview, GPT-5.4/5.4 Pro, and Gemini 3.1 Pro.
Original inputs and harness: Vellum did not publish the complete inputs, code, or run configuration used to rerun these benchmarks; this is an analysis of the data and of task selection.
| Benchmark | Opus 4.7 | Opus 4.6 | Key comparison |
|---|---|---|---|
| SWE-bench Verified | 87.6% | 80.8% | Gemini 3.1 Pro 80.6% |
| SWE-bench Pro | 64.3% | 53.4% | GPT-5.4 57.7%; Gemini 3.1 Pro 54.2% |
| Terminal-Bench 2.0 | 69.4% | 65.4% | GPT-5.4 75.1%; Gemini 3.1 Pro 68.5% |
| MCP-Atlas | 77.3% | 75.8% | GPT-5.4 68.1%; Gemini 3.1 Pro 73.9% |
| Finance Agent v1.1 | 64.4% | 60.1% | GPT-5.4 Pro 61.5%; Gemini 3.1 Pro 59.7% |
| OSWorld-Verified | 78.0% | 72.7% | GPT-5.4 75.0% |
| BrowseComp | 79.3% | 83.7% | GPT-5.4 Pro 89.3%; Gemini 3.1 Pro 85.9% |
| GPQA Diamond | 94.2% | 91.3% | Gemini 3.1 Pro 94.3% |
| CharXiv without tools/with tools | 82.1% / 91.0% | 69.1% / 84.7% | Mythos Preview 86.1% / 93.2% |
Vellum also records the official migration note: the 4.7 tokenizer may turn the same text into approximately 1.0–1.35× as many tokens; long Agent turns at high effort may also increase output tokens.
If a workflow's bottleneck is complex coding, tool calls, screen understanding, or specialized analysis, Opus 4.7's gains are worth an A/B migration test. If the bottleneck is deep web search, its BrowseComp score of 79.3% is below 4.6's 83.7%, so consider retaining 4.6 or comparing other models.
The article is primarily a second-order interpretation of official results, not an independent benchmark rerun by Vellum using the same harness.
Official partner figures (CursorBench, 93-task coding, visual acuity, and so on) lack the complete task definitions and raw outputs, so they should not be given equal weight with the public data.
The selection conclusions depend on the specific tools, context, effort, and token budget; the article does not provide a standardized API parameter table.
Choose business samples from each of the four task categories listed in the article: repository issues, multi-tool workflows, visual charts, and web research.
Hold the prompt, tool schema, permissions, effort, and context limit constant, then run Opus 4.6 and 4.7 separately.
Record task success, tests passed, tool errors, tokens, latency, factual errors, and the amount of human revision.
Examine BrowseComp-style research tasks separately so that coding scores do not average away a regression in web research.
Estimate the per-task cost after the tokenizer change using real traffic, then decide whether to upgrade or route tasks across models.
The Vellum article explains SWE-bench Pro, MCP-Atlas, OSWorld, BrowseComp, GPQA, HLE, and CharXiv item by item, and explicitly notes that BrowseComp fell from 83.7% to 79.3%; the remaining table data comes from Anthropic's reported system card results.
Vellum's overall assessment is “a focused, high-performance upgrade,” rather than a sweeping win across the board; this is consistent with the split between the gains in coding/tool use and the decline in BrowseComp shown in the table.
Claude Opus 4.7