Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

Claude Opus 4.8 · Media / benchmark · Independent measurement

Claude Opus 4.8: Vellum's Cross-Model Benchmark Comparison and Harness Boundaries

Vellum’s comparison frames Claude Opus 4.8 across named benchmark tasks and model choices; mixed harnesses prevent a single cross-task ranking.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Condition
Model/version: claude-opus-4-8; use Vellum’s table snapshot.
Condition
Harness/tasks: cross-model tables mix tasks, settings, and clients; raw traces are not fully public.
Condition
Sample/date: interpret the article snapshot; reopened 2026-09-20.

Key data and applicable tasks

One-sentence takeaway

Vellum's line-by-line compilation of Anthropic's system card shows Opus 4.8 leading in most public comparisons, but the harness differences in Terminal-Bench and the counterexample from Finance Agent v2 show that model selection must be retested in the same tool environment and on the relevant vertical task.

Use cases

  • Suitable tasks: Comparing public signals for coding, terminal agents, reasoning, computer use, and professional knowledge work.

  • Unsuitable tasks: Treating the article's secondary compilation as your own production acceptance test, or extrapolating a single benchmark directly to every domain.

  • Applicable model versions: Claude Opus 4.8, compared with Opus 4.7, GPT-5.5, and Gemini 3.1 Pro; Finance Agent v2 also includes Gemini 3.5 Flash.

  • Applicable client, agent, or API: The article discusses harnesses such as Terminus-2/Codex CLI in the system card; rerun the tests with your own tool stack.

  • Recommended reasoning tier and parameters: The article mainly relays the official default effort; full run parameters, random seeds, and inputs were not disclosed, so do not use this as the basis for fixing parameters.

Test environment

  • Data/method: Vellum compiled the public results in Table 8.1.A of the Anthropic Claude Opus 4.8 System Card item by item and explained the impact of different harnesses.

  • Comparison models: Opus 4.7, GPT-5.5, and Gemini 3.1 Pro; Finance Agent v2 also includes Gemini 3.5 Flash.

  • Reproducibility: The source provides benchmark names, scores, and some harness details, but does not publish the original inputs, complete configurations, or independent run logs. Therefore, “reproducible” here is limited to rechecking the public tables and benchmark names; it cannot be described as an independent rerun by Vellum.

Inputs/configuration

  • SWE-Bench Pro: Multi-file issues in active repositories; the article notes that it is harder than ordinary SWE-bench and less susceptible to memorization leakage.

  • Terminal-Bench 2.1: Compared under the same public Terminus-2 harness; GPT-5.5 also has a Codex CLI headline, and its number cannot be directly compared with the Terminus-2 figures.

  • OSWorld-Verified: Document editing, web browsing, and file management tasks in a real Ubuntu VM.

  • HLE, GDPval-AA, Finance Agent v2: The article reports scores and task types from the public system cards, but does not provide per-question inputs or run logs.

Results data

BenchmarkOpus 4.8Main comparisons and notes
SWE-Bench Pro pass rate69.2%Opus 4.7 64.3%, GPT-5.5 58.6%, Gemini 3.1 Pro 54.2%
SWE-bench Verified88.6%Opus 4.7 87.6%, Gemini 3.1 Pro 80.6%
Terminal-Bench 2.1 / Terminus-274.6%GPT-5.5 Terminus-2 78.2%, Gemini 3.1 Pro 70.3%, Opus 4.7 66.1%; GPT-5.5's Codex CLI headline is 83.4%
HLE (without tools)49.8%Opus 4.7 46.9%, Gemini 3.1 Pro 44.4%, GPT-5.5 41.4%
HLE (with tools)57.9%Opus 4.7 54.7%, GPT-5.5 52.2%, Gemini 3.1 Pro 51.4%
OSWorld-Verified pass@183.4%Opus 4.7 82.8%, GPT-5.5 78.7%, Gemini 3.1 Pro 76.2%
GDPval-AA1,890GPT-5.5 1,769, Opus 4.7 1,753, Gemini 3.1 Pro 1,314
Finance Agent v253.9%Gemini 3.5 Flash 57.9%, GPT-5.5 51.8%, Opus 4.7 51.5%, Gemini 3.1 Pro 43.0%

Conclusion

In the public comparisons, Opus 4.8 leads on tasks including SWE-Bench Pro, HLE, OSWorld-Verified, and GDPval-AA. However, GPT-5.5's Terminal-Bench result changes from 83.4% (Codex CLI) to 78.2% (Terminus-2) with the harness, while Finance Agent v2 is led by the smaller, faster Gemini 3.5 Flash. The most defensible conclusion is that Opus 4.8 is well suited to complex coding and knowledge work, but vertical domains and tool environments still require separate evaluation.

Limitations

  • Vellum's figures mainly come from Anthropic's system card rather than original experiments rerun by Vellum; the article does not include complete inputs, configurations, random seeds, or original logs.

  • System card scores are affected by harnesses, tools, token limits, and version revisions; the Opus 4.7 figure for OSWorld also notes a zoom-tool fix and a change to the max-token limit.

  • The task distributions, scoring criteria, and costs for HLE, GDPval-AA, and Finance Agent v2 are not disclosed in full in the article; they cannot be used to estimate real-world ROI.

  • The article's “leading in most” does not mean leading in every domain, and the counterexample from Finance Agent v2 must not be overlooked.

Reproduction steps

  1. First fix the same model version, tool harness, container/VM, token limit, and scoring script.

  2. Prioritize rerunning SWE-Bench Pro, Terminal-Bench 2.1, and OSWorld-Verified, while recording pass rate, tool steps, tokens, duration, and failure types.

  3. Run models such as GPT-5.5 separately on Codex CLI and Terminus-2; do not place figures from different harnesses in the same column for direct comparison.

  4. Build a separate Finance Agent or knowledge-work subset for the target business, report confidence intervals and cost, and do not treat public scores as a production guarantee.

Source excerpt or observation (short quote for compliance only)

Vellum's most valuable observation is not “how many categories Opus 4.8 won,” but that “harness matters as much as the model” in Terminal-Bench. This is the applicability boundary that must be retained when reviewing this comparison.

What this supports

  • Supports selection discussion within Vellum’s published tables and task categories.

What this does not support

  • Does not support cross-harness ranking or a stable current leaderboard.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Vellum · Nicolas Zeeb · Original publication date 2026-05-28 · Site edit date 2026-09-20

Open original source

Claude Opus 4.8

Compare Claude Opus 4.8 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

Claude Opus 4.8: What It Is, What Changed, and Its Legacy Status

A sourced Claude Opus 4.8 overview covering effort, context, pricing, provider access, the 4.7 upgrade, safety limits and its current legacy lifecycle.

Related reviews

Claude Opus 4.8: Official Release Capabilities, Agent Workflows, and Honesty BoundariesAnthropic’s Opus 4.8 release highlights more reliable agent judgment, effort control, and dynamic workflows; tester quotes and vendor evals do not establish independent production success.Claude Opus 4.8: Effort Levels, Tool Use, and Code Review Prompt PatternsFollow a task-specific guide for “Claude Opus 4.8: Effort Levels, Tool Use, and Code Review Prompt Patterns”; prerequisites, steps, checks, fixes, and source boundaries are explicit.Claude Opus 4.8: Max-Effort Prompt for High-Stakes TasksFollow a task-specific guide for “Claude Opus 4.8: Max-Effort Prompt for High-Stakes Tasks”; prerequisites, steps, checks, fixes, and source boundaries are explicit.