Anthropic positions Opus 4.8 as a steady upgrade for long-horizon coding, Agent workflows, and professional work: it defaults to high effort, supports xhigh/max, and is more reliable with tools and long-running tasks, but price, token usage, tool harnesses, and safety boundaries should all be reviewed.
Suitable tasks: Long-horizon coding, browser/computer Agents, professional knowledge work, complex analysis, and workflows that need to recognize their own uncertainty.
Unsuitable tasks: Model selection based only on a single leaderboard score, high-risk automated decisions without a verification step, and any task where cost is highly sensitive.
Applicable model version: Claude Opus 4.8, API ID claude-opus-4-8.
Applicable clients, Agents, or APIs: Claude, Claude Code, Cowork, and the Claude API; the dynamic workflow is a research-preview feature for Claude Code Enterprise, Team, and Max.
Recommended reasoning levels and parameters: The official default is high; extra/xhigh are recommended for difficult tasks and long-running asynchronous workflows, while the highest max level requires weighing token usage and possible overthinking. Standard pricing is $5 per million input tokens and $25 per million output tokens; fast mode costs $10/$50.
Model/version: Claude Opus 4.8; the official release page compares it with Opus 4.7 and other models and links to the system card.
Tools/clients: Claude Code, Claude API, and client-side effort controls; dynamic workflows can run many sub-Agents in parallel.
Evaluation sources: Capability evaluations published by Anthropic, the system card, and feedback from early testers; the release page does not provide all test inputs, random seeds, or the complete configuration for each item.
The release page says Opus 4.8 defaults to high effort; users can select xhigh or max.
The Messages API adds the ability to insert a system entry in the messages array, allowing permissions, token budgets, or environment context to be updated during a task without going through a user turn.
Dynamic workflows can plan the work, run hundreds of sub-Agents in parallel, and validate outputs before reporting; the official example is a migration at the scale of hundreds of thousands of lines of code.
The official release page says Opus 4.8 scores 84% on Online-Mind2Web and describes this as a significant improvement in computer-use and browser-Agent capabilities.
The official evaluation says it is about four times less likely than its predecessor to let code it wrote “pass without being prompted” about defects; this describes the relative risk of errors going unnoticed and does not mean the error rate is fixed at a fourfold reduction across all projects.
The release page cites early testers who completed every case on the Super-Agent benchmark, used fewer tool-calling steps on CursorBench, and exceeded 10% overall all-pass for the first time on the Legal Agent Benchmark; these are partner/customer reports, and the complete experimental configurations are not disclosed on that page.
Official pricing: standard mode costs $5 per million input tokens and $25 per million output tokens; fast mode costs $10 per million input tokens and $50 per million output tokens, with speed of about 2.5×.
The practical value of Opus 4.8 lies not only in leaderboard scores, but also in self-questioning during long-horizon tasks, tool-calling efficiency, context coordination, and adjustable effort. Real-world systems should separate “identify problems—verify—execute” and use task sets to measure tokens, latency, tool steps, and failure recovery, rather than merely repeating official marketing claims.
This is material released by the model vendor; some results come from Anthropic’s internal evaluations or are self-reported by partners, and cannot replace independent reproduction.
The release page does not show the complete inputs, sample sizes, confidence intervals, operating costs, or harnesses for all benchmarks; detailed figures should be checked against the linked system card and independent comparisons.
“Four times less likely to let defects go unreported” is a relative conclusion from the official evaluation and cannot be generalized into a universal guarantee of production quality.
The availability, quotas, and prices of dynamic workflows, fast mode, and effort levels may change by client or plan; verify them again before launch.
Run a set of long-horizon coding tasks with Opus 4.8 and Opus 4.7 using the same code repository, tool definitions, context, and max_tokens.
Test high, xhigh, and max separately, recording success rate, tokens, total duration, number of tool calls, number of manual takeovers, and number of undetected defects.
Using the same test set and review criteria, separate “discovery” from “verification” and record recall, precision, and F1.
For browser Agents, record pass@1 on Online-Mind2Web-like tasks, recovery failures, and cost per task; do not mix leaderboard figures from different harnesses.
The official page describes Opus 4.8 as a “modest but tangible improvement,” while presenting default high effort, dynamic workflows, effort controls, and the Messages API system entry as accompanying changes. Taken together, this indicates that deployment configuration and workflow design can materially affect results.
Claude Opus 4.8