Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
MediaClaude Fable 5.1

Artificial Analysis: Claude Fable 5.1 Intelligence Index, Task Breakdowns, and Cost

Original source

Artificial Analysis

AuthorArtificial Analysis

Source date2026-09-01

Tabbit curation2026-09-08

Read original

One-sentence takeaway

Artificial Analysis's independent pre-release evaluation shows Fable 5.1 reaching 66 on the Intelligence Index and setting records across multiple agentic tasks, but its per-task cost at max is higher than Fable 5's; xhigh is often the more balanced cost/capability point.

Use cases

  • Suitable tasks: High-difficulty knowledge work, agentic coding, science and math problems, financial tool calling, and workflows that require high-quality professional deliverables.

  • Unsuitable tasks: High-concurrency simple tasks that are extremely sensitive to cost and output tokens; presentation tasks with extremely high standards for visual quality, where Fable 5.1 trails Opus 5 on the presentation subscore of AA-Briefcase.

  • Applicable model versions: Claude Fable 5.1, compared at Low/Medium/High/xhigh/max effort; some requests used Anthropic's default server-side fallback.

  • Applicable client, Agent, or API: Artificial Analysis's open reference Agent harness, Stirrup; these are not end-to-end product scores for Claude.ai or Claude Code.

  • Recommended reasoning tier and parameters: First compare xhigh and max as the quality ceiling; when costs need to be reduced, test xhigh/medium rather than assuming max is always optimal.

Test method/data

Artificial Analysis says it evaluated Fable 5.1 during the pre-release phase and used Anthropic's default server-side fallback; when requests triggered safety guardrails, about 4% of the Intelligence Index output tokens were produced by Opus 4.8 or Opus 5 rather than Fable 5.1. This boundary must be recorded alongside the results.

Core results:

  • Intelligence Index: Fable 5.1 max scored 66, higher than Opus 5 max at 63, Fable 5 max at 62, GPT-5.6 Sol max at 61, and Grok 4.6 high at 61; Fable 5.1 improved by 4 points relative to Fable 5.

  • Humanity's Last Exam: 59.1%, higher than Fable 5's previous 55.5%.

  • Terminal-Bench v2.1: 91.4%; SciCode: 62.0%; the article says both were the highest scores it had measured at that time.

  • 𝜏³-Banking: up 9 percentage points relative to Fable 5.

  • GDPval-AA v2: Fable 5.1 max 1853 Elo, versus Fable 5 max at 1723; Opus 5 max was 1824, with the difference within the confidence interval.

  • AA-Briefcase: Fable 5.1 max 1694 Elo, versus Opus 5 max at 1685; the article judges them essentially tied. Fable 5.1's analytical quality was 2025 and its presentation score was 1495, while Opus 5 scored 1980 and 1572 respectively.

  • AA-Omniscience: Fable 5.1 max attempted to answer 93.4% of questions, with an accuracy rate of 67.2%; Fable 5 reached 87.8% and 65.4%, respectively. Among questions it did not answer correctly, Fable 5.1's attempt rate was 72.6%, higher than Fable 5's 63.6%, so the overall index was tied with Fable 5.

  • Cost: Fable 5.1 max costs about $3.76 per Intelligence Index task, compared with $3.14 for Fable 5 max and $2.34 for Opus 5 max; Fable 5.1 used about 1.7 times Fable 5's output tokens. xhigh scored 65 at about $2.72 per task.

  • Effort range: Fable 5.1's output tokens increased from about 13.1M at low to 143.7M at max, while the Intelligence Index rose from 58 to 66, a span of about 11 times.

Conclusions

  • If the goal is high-difficulty research, agentic coding, or tool-oriented knowledge work, Fable 5.1's strengths are high-end capability and task completion quality, not simply a low price.

  • If a slightly lower top score is acceptable, xhigh's 65 points/$2.72 is more suitable as a cost-sensitive default tier than max's 66 points/$3.76; it still needs to be validated on your own task set.

  • GDPval-AA v2 and AA-Briefcase indicate that it is strong in analytical quality, but the “presentation quality of professional deliverables” does not necessarily lead at the same time; sub-scores should be reported separately.

Boundaries and risks

  • The evaluation provider says it used the reference Agent harness and default fallback; about 4% of output tokens came from an Opus fallback, so the results cannot simply be described as “100% native Fable 5.1 results.”

  • The Intelligence Index aggregates multiple sub-evaluations and cannot replace success rates for specific tasks; the models also differ in effort, output tokens, and tool trajectories.

  • The gap between Fable 5.1 and Opus 5 on GDPval-AA v2 is within the confidence interval; 1853 should not be interpreted as a statistically significant overall lead.

  • AA-Omniscience shows that a higher attempt rate also brings more incorrect attempts; it should be broken down into “correct-answer rate, refusal/non-attempt rate, and hallucination rate,” rather than looking only at accuracy.

  • The article is an aggregated report published by the evaluation provider; complete per-question raw outputs and all parameters require further verification against Artificial Analysis's model pages/data products.

Reproduction recommendations

  1. Fix Fable 5.1's effort, fallback strategy, and tool harness, and explicitly state whether routing to Opus after a safety trigger is allowed.

  2. Reproduce the same task set, recording completion rate, rubric score, output tokens, cache hits, per-task cost, and fallback ratio separately.

  3. For GDPval/Briefcase-style professional deliverables, analyze analytical quality and presentation quality separately to avoid letting a single aggregate score conceal differences.

  4. Run paired repeated tests with xhigh, max, Fable 5, and Opus 5, and report confidence intervals or standard errors.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Fable 5.1

Use and compare models in Tabbit

Claude Fable 5.1

Related reviews

OfficialOpenRouter Model Page2026-09-02

OpenRouter: Provider Performance and Benchmark Snapshot for Claude Fable 5.1

MediaAnthropic News / Introducing Claude Fable 5.1 and Claude Mythos 5.12026-09

Claude Fable 5.1 Official Release: Multiple Benchmarks, Cost Tiers, and Safety Boundaries

MediaSimpleBench

SimpleBench: Claude Fable 5.1's Everyday Reasoning and Human Baseline

CommunityReddit / r/ArtificialInteligence2026-09-02

Reddit community: Skepticism about Fable 5.1 benchmarks and task-tiered experience

Claude Fable 5.1

Related prompts

MediaAnthropic Claude Platform Docs

Claude Fable 5.1 Official Prompting Methods: Writing Density, Batched Tool Calls, and Task Completion

CommunityReddit (r/claude)2026-09-02

Reddit Prompting Techniques: Long-form Analysis, Counterarguments, and Concise Execution

CommunityReddit (r/ClaudeAI)2026-09-05

Reddit case: Fable 5.1 + Blender MCP generates a large world region in one shot