Artificial Analysis's independent pre-release evaluation shows Fable 5.1 reaching 66 on the Intelligence Index and setting records across multiple agentic tasks, but its per-task cost at max is higher than Fable 5's; xhigh is often the more balanced cost/capability point.
Suitable tasks: High-difficulty knowledge work, agentic coding, science and math problems, financial tool calling, and workflows that require high-quality professional deliverables.
Unsuitable tasks: High-concurrency simple tasks that are extremely sensitive to cost and output tokens; presentation tasks with extremely high standards for visual quality, where Fable 5.1 trails Opus 5 on the presentation subscore of AA-Briefcase.
Applicable model versions: Claude Fable 5.1, compared at Low/Medium/High/xhigh/max effort; some requests used Anthropic's default server-side fallback.
Applicable client, Agent, or API: Artificial Analysis's open reference Agent harness, Stirrup; these are not end-to-end product scores for Claude.ai or Claude Code.
Recommended reasoning tier and parameters: First compare xhigh and max as the quality ceiling; when costs need to be reduced, test xhigh/medium rather than assuming max is always optimal.
Artificial Analysis says it evaluated Fable 5.1 during the pre-release phase and used Anthropic's default server-side fallback; when requests triggered safety guardrails, about 4% of the Intelligence Index output tokens were produced by Opus 4.8 or Opus 5 rather than Fable 5.1. This boundary must be recorded alongside the results.
Core results:
Intelligence Index: Fable 5.1 max scored 66, higher than Opus 5 max at 63, Fable 5 max at 62, GPT-5.6 Sol max at 61, and Grok 4.6 high at 61; Fable 5.1 improved by 4 points relative to Fable 5.
Humanity's Last Exam: 59.1%, higher than Fable 5's previous 55.5%.
Terminal-Bench v2.1: 91.4%; SciCode: 62.0%; the article says both were the highest scores it had measured at that time.
𝜏³-Banking: up 9 percentage points relative to Fable 5.
GDPval-AA v2: Fable 5.1 max 1853 Elo, versus Fable 5 max at 1723; Opus 5 max was 1824, with the difference within the confidence interval.
AA-Briefcase: Fable 5.1 max 1694 Elo, versus Opus 5 max at 1685; the article judges them essentially tied. Fable 5.1's analytical quality was 2025 and its presentation score was 1495, while Opus 5 scored 1980 and 1572 respectively.
AA-Omniscience: Fable 5.1 max attempted to answer 93.4% of questions, with an accuracy rate of 67.2%; Fable 5 reached 87.8% and 65.4%, respectively. Among questions it did not answer correctly, Fable 5.1's attempt rate was 72.6%, higher than Fable 5's 63.6%, so the overall index was tied with Fable 5.
Cost: Fable 5.1 max costs about $3.76 per Intelligence Index task, compared with $3.14 for Fable 5 max and $2.34 for Opus 5 max; Fable 5.1 used about 1.7 times Fable 5's output tokens. xhigh scored 65 at about $2.72 per task.
Effort range: Fable 5.1's output tokens increased from about 13.1M at low to 143.7M at max, while the Intelligence Index rose from 58 to 66, a span of about 11 times.
If the goal is high-difficulty research, agentic coding, or tool-oriented knowledge work, Fable 5.1's strengths are high-end capability and task completion quality, not simply a low price.
If a slightly lower top score is acceptable, xhigh's 65 points/$2.72 is more suitable as a cost-sensitive default tier than max's 66 points/$3.76; it still needs to be validated on your own task set.
GDPval-AA v2 and AA-Briefcase indicate that it is strong in analytical quality, but the “presentation quality of professional deliverables” does not necessarily lead at the same time; sub-scores should be reported separately.
The evaluation provider says it used the reference Agent harness and default fallback; about 4% of output tokens came from an Opus fallback, so the results cannot simply be described as “100% native Fable 5.1 results.”
The Intelligence Index aggregates multiple sub-evaluations and cannot replace success rates for specific tasks; the models also differ in effort, output tokens, and tool trajectories.
The gap between Fable 5.1 and Opus 5 on GDPval-AA v2 is within the confidence interval; 1853 should not be interpreted as a statistically significant overall lead.
AA-Omniscience shows that a higher attempt rate also brings more incorrect attempts; it should be broken down into “correct-answer rate, refusal/non-attempt rate, and hallucination rate,” rather than looking only at accuracy.
The article is an aggregated report published by the evaluation provider; complete per-question raw outputs and all parameters require further verification against Artificial Analysis's model pages/data products.
Fix Fable 5.1's effort, fallback strategy, and tool harness, and explicitly state whether routing to Opus after a safety trigger is allowed.
Reproduce the same task set, recording completion rate, rubric score, output tokens, cache hits, per-task cost, and fallback ratio separately.
For GDPval/Briefcase-style professional deliverables, analyze analytical quality and presentation quality separately to avoid letting a single aggregate score conceal differences.
Run paired repeated tests with xhigh, max, Fable 5, and Opus 5, and report confidence intervals or standard errors.
Claude Fable 5.1