Vellum's line-by-line compilation of Anthropic's system card shows Opus 4.8 leading in most public comparisons, but the harness differences in Terminal-Bench and the counterexample from Finance Agent v2 show that model selection must be retested in the same tool environment and on the relevant vertical task.
Suitable tasks: Comparing public signals for coding, terminal agents, reasoning, computer use, and professional knowledge work.
Unsuitable tasks: Treating the article's secondary compilation as your own production acceptance test, or extrapolating a single benchmark directly to every domain.
Applicable model versions: Claude Opus 4.8, compared with Opus 4.7, GPT-5.5, and Gemini 3.1 Pro; Finance Agent v2 also includes Gemini 3.5 Flash.
Applicable client, agent, or API: The article discusses harnesses such as Terminus-2/Codex CLI in the system card; rerun the tests with your own tool stack.
Recommended reasoning tier and parameters: The article mainly relays the official default effort; full run parameters, random seeds, and inputs were not disclosed, so do not use this as the basis for fixing parameters.
Data/method: Vellum compiled the public results in Table 8.1.A of the Anthropic Claude Opus 4.8 System Card item by item and explained the impact of different harnesses.
Comparison models: Opus 4.7, GPT-5.5, and Gemini 3.1 Pro; Finance Agent v2 also includes Gemini 3.5 Flash.
Reproducibility: The source provides benchmark names, scores, and some harness details, but does not publish the original inputs, complete configurations, or independent run logs. Therefore, “reproducible” here is limited to rechecking the public tables and benchmark names; it cannot be described as an independent rerun by Vellum.
SWE-Bench Pro: Multi-file issues in active repositories; the article notes that it is harder than ordinary SWE-bench and less susceptible to memorization leakage.
Terminal-Bench 2.1: Compared under the same public Terminus-2 harness; GPT-5.5 also has a Codex CLI headline, and its number cannot be directly compared with the Terminus-2 figures.
OSWorld-Verified: Document editing, web browsing, and file management tasks in a real Ubuntu VM.
HLE, GDPval-AA, Finance Agent v2: The article reports scores and task types from the public system cards, but does not provide per-question inputs or run logs.
| Benchmark | Opus 4.8 | Main comparisons and notes |
|---|---|---|
| SWE-Bench Pro pass rate | 69.2% | Opus 4.7 64.3%, GPT-5.5 58.6%, Gemini 3.1 Pro 54.2% |
| SWE-bench Verified | 88.6% | Opus 4.7 87.6%, Gemini 3.1 Pro 80.6% |
| Terminal-Bench 2.1 / Terminus-2 | 74.6% | GPT-5.5 Terminus-2 78.2%, Gemini 3.1 Pro 70.3%, Opus 4.7 66.1%; GPT-5.5's Codex CLI headline is 83.4% |
| HLE (without tools) | 49.8% | Opus 4.7 46.9%, Gemini 3.1 Pro 44.4%, GPT-5.5 41.4% |
| HLE (with tools) | 57.9% | Opus 4.7 54.7%, GPT-5.5 52.2%, Gemini 3.1 Pro 51.4% |
| OSWorld-Verified pass@1 | 83.4% | Opus 4.7 82.8%, GPT-5.5 78.7%, Gemini 3.1 Pro 76.2% |
| GDPval-AA | 1,890 | GPT-5.5 1,769, Opus 4.7 1,753, Gemini 3.1 Pro 1,314 |
| Finance Agent v2 | 53.9% | Gemini 3.5 Flash 57.9%, GPT-5.5 51.8%, Opus 4.7 51.5%, Gemini 3.1 Pro 43.0% |
In the public comparisons, Opus 4.8 leads on tasks including SWE-Bench Pro, HLE, OSWorld-Verified, and GDPval-AA. However, GPT-5.5's Terminal-Bench result changes from 83.4% (Codex CLI) to 78.2% (Terminus-2) with the harness, while Finance Agent v2 is led by the smaller, faster Gemini 3.5 Flash. The most defensible conclusion is that Opus 4.8 is well suited to complex coding and knowledge work, but vertical domains and tool environments still require separate evaluation.
Vellum's figures mainly come from Anthropic's system card rather than original experiments rerun by Vellum; the article does not include complete inputs, configurations, random seeds, or original logs.
System card scores are affected by harnesses, tools, token limits, and version revisions; the Opus 4.7 figure for OSWorld also notes a zoom-tool fix and a change to the max-token limit.
The task distributions, scoring criteria, and costs for HLE, GDPval-AA, and Finance Agent v2 are not disclosed in full in the article; they cannot be used to estimate real-world ROI.
The article's “leading in most” does not mean leading in every domain, and the counterexample from Finance Agent v2 must not be overlooked.
First fix the same model version, tool harness, container/VM, token limit, and scoring script.
Prioritize rerunning SWE-Bench Pro, Terminal-Bench 2.1, and OSWorld-Verified, while recording pass rate, tool steps, tokens, duration, and failure types.
Run models such as GPT-5.5 separately on Codex CLI and Terminus-2; do not place figures from different harnesses in the same column for direct comparison.
Build a separate Finance Agent or knowledge-work subset for the target business, report confidence intervals and cost, and do not treat public scores as a production guarantee.
Vellum's most valuable observation is not “how many categories Opus 4.8 won,” but that “harness matters as much as the model” in Terminal-Bench. This is the applicability boundary that must be retained when reviewing this comparison.
Claude Opus 4.8