Anthropic's release data positions Fable 5.1 as a high-end agent model for long-horizon coding, scientific research, and knowledge work, but different effort levels, tools, harnesses, and production safety guardrails can materially change the results, so a single leaderboard number is not enough.
Suitable tasks: Long-horizon coding and debugging, scientific-research agents, knowledge work, desktop/browser operation, complex multi-step analysis, and workflows that require continuous verification.
Unsuitable tasks: Simple questions that only require a low-latency short answer and have no tools or verification step; for high-risk dual-use cybersecurity and life-science tasks, Fable's guardrails may hand off to another model or refuse.
Applicable model version: Claude Fable 5.1; the page also lists Fable 5, Opus 5, and GPT-5.6 Sol for comparison.
Applicable client, agent, or API: The release page does not limit usage to a single client; the charts and benchmarks use their respective harnesses and should not be assumed to match results from Claude.ai, Claude Code, or a third-party API.
Recommended reasoning tiers and parameters: Anthropic recommends starting with High, then sweeping Low, Medium, xhigh, and max on your own evals; the release page says Claude Code defaults to High, while Claude Cowork and Claude.ai default to Medium.
The release page does not disclose the complete questions, prompts, repeated-sampling counts, or all harness configurations for each benchmark, so the figures below are official release baselines, not independently reproduced results.
| Task/benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | Notes |
|---|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% | Agentic scientific research; the official chart notes a standard error of approximately ±3.5–4.5 points |
| Terminal-Bench 4.0 | 55.8% | 52.3% | 42.0% | 37.3% | Agentic coding; the page separately lists Mythos 5.1 at 60.9%, which must not be mixed with Fable's guardrail conditions |
| GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 | Knowledge work |
| OSWorld 2.0 (partial) | 77.9% | 72.9% | 75.4% | — | Computer use; GPT-5.6 Sol was not provided |
| OSWorld 2.0 (strict) | 41.7% | 36.1% | 39.6% | — | Computer use; this uses a different definition from partial |
| Humanity’s Last Exam (no tools) | 60.9% | 57.8% | 56.6% | — | Multidisciplinary reasoning |
| Humanity’s Last Exam (with tools) | 65.0% | 63.8% | 63.6% | — | The tool condition changes, so this cannot be directly compared with the no-tools result |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% | Business workflows |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% | Agentic coding |
The release page also provides a verification note for Terminal-Bench-Science: the public leaderboard uses the Claude Code harness and three trials per question, with Opus 5 at 30.0% and Fable 5 at 21.4%; Anthropic's own setup reproduced 29.0% and 24.7%, respectively, and says both are within the noise range. This shows that harness differences alone can materially change the score.
Scientific research and root-cause analysis are the strongest candidates for distinctive strengths: Terminal-Bench-Science is 52.6%, higher than the official comparison models; however, the standard error and experimental harness must be retained, and the result should not be interpreted as roughly doubling performance on all scientific tasks.
Knowledge work and engineering agents are in the frontier range: GDPval-AA v2 is 1853 and CursorBench is 73.4%, making the model suitable for cross-file changes, complex research, and long-running workflows with verification.
Route by effort for the cost/quality trade-off: Anthropic says Fable 5.1 can approach or exceed Fable 5's results at Low/Medium; actual selection should record quality, tokens, latency, and cost per task rather than comparing max scores alone.
These are self-reported results from the model provider; the page does not disclose all original questions, per-question traces, prompts, variance, or complete API parameters.
Anthropic acknowledges that production safety guardrails affect benchmarks: on OSWorld tasks where guardrails intervened, Fable 5.1 and Fable 5 were recorded as zero; AutomationBench had a similar case for Fable 5. Other restricted cybersecurity and life-science tasks were completed by Opus models, so not every number in the table should be treated as “pure Fable 5.1 capability.”
The public Terminal-Bench-Science leaderboard and the comparison scores from Anthropic's setup differ; reproduction must fix the question-set version, harness, effort, tools, and safety settings.
Partner quotes in the release page (such as Millennium, MongoDB, and Browserbase) are customer case studies. The complete test sets and methods were not published with the page, so they cannot be interpreted as independent controlled benchmarks.
Fable 5.1 and Mythos 5.1 use the same model but different guardrails; the Mythos scores on the page cannot be directly treated as results from publicly callable Fable.
Fix the claude-fable-5-1 model snapshot, effort (at least Low/Medium/High/max), maximum output, tool definitions, and safety routing.
Build separate task sets for coding, scientific retrieval, knowledge work, and computer use, each with explicit success criteria; record the complete inputs, tool traces, output tokens, latency, cost, and guardrail interventions.
Repeat each task at least 3 times and report the mean, standard error, and failure type; do not substitute a single run for the official multi-trial comparison.
Treat the table on this page as an “official baseline,” and separately compare it with independent sources such as Artificial Analysis and SimpleBench.
Claude Fable 5.1