Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
Review
OfficialClaude Fable 5.1

OpenRouter: Provider Performance and Benchmark Snapshot for Claude Fable 5.1

Original source

OpenRouter Model Page

AuthorOpenRouter

Source date2026-09-02

Tabbit curation2026-09-08

Read original

One-sentence takeaway

OpenRouter's live page shows Fable 5.1 delivering roughly 44–56 tok/s in one-week average throughput across multiple Providers, with roughly 4.18–6.01 seconds of average first-chunk/request latency, and provides routing snapshots for GPQA Diamond, TAU-Bench, and structured-output error rates. It is suitable for choosing an integration layer, not for replacing controlled capability evaluations.

Use cases

  • Suitable tasks: Comparing latency, throughput, availability, and structured-output performance when connecting to the same model through Azure, Anthropic, Amazon Bedrock BYOK, and Google Vertex.

  • Unsuitable tasks: Inferring the model's reasoning ability, long-horizon Agent success rates, or the end-to-end experience across different clients; Provider routing and traffic composition affect the results.

  • Applicable model version: anthropic/claude-fable-5.1 on OpenRouter.

  • Applicable clients, Agents, or APIs: OpenRouter routing; the page lists high-traffic applications such as Hermes Agent, Claude Code, and Kilo Code, but this is not a controlled comparison experiment for those applications.

  • Recommended reasoning tier and parameters: The page does not disclose the effort, prompt, output length, or sampling settings corresponding to this telemetry; production integrations should fix these variables independently.

Test method/data

The page displays P50 runtime data for each Provider:

  • Standard endpoint pricing is $10/million tokens for input, $50/million tokens for output, and $0.25/million tokens for cache reads.

  • Provider P50: Azure latency 14.96s and throughput 56 tok/s; Anthropic latency 5.97s, throughput 45 tok/s, and availability 99.91%; Amazon Bedrock (BYOK) latency 11.58s and throughput 53 tok/s; Google Vertex latency 4.84s, throughput 65 tok/s, and availability 99.35%.

  • Average throughput over the past week: Vertex 56 tok/s, Bedrock 46 tok/s, and Anthropic 44 tok/s; average first-chunk/request latency: Vertex 4.18s, Anthropic 5.96s, and Azure 6.01s; end-to-end average latency was 12.51s, 16.20s, and 16.15s, respectively.

  • AutoExacto Benchmarks: GPQA Diamond—automatic routing 90.9%, Anthropic 86.5%, Vertex 84.9%, and Azure 86.3%; TAU-Bench—Anthropic 79.3%, Vertex 76.7%, and Azure 74.7%. The page does not provide the complete question sets or repetition counts for the automatic-routing and Provider results.

  • Structured-output error rates: Anthropic averaged 8.33%, Bedrock 33.12%, and Azure 38.51%; cache hit rates were approximately 85.26%–86.82%.

  • At collection time, the page showed OpenRouter uptime of 100% and availability of 99.92% over the past 3 days, while "no-routing" availability, which does not bypass Provider failures, was 98.25%.

Conclusion

OpenRouter's evidence is primarily useful for deployment-layer decisions: Vertex led in throughput and latency in this page snapshot, while the Anthropic endpoint had a lower structured-output error rate; if a task depends on JSON/structured calls, Provider differences may affect stability more than the model name itself. In production, fix the provider filter, routing strategy, and retry method before measuring again.

Limitations and risks

  • This is platform live telemetry, not a public controlled benchmark report; the page does not provide the complete input set, output-length distribution, effort, sampling parameters, or statistical confidence intervals.

  • P50, past-week averages, and past-three-days uptime use different time windows and cannot be treated as one unified metric.

  • OpenRouter automatically switches when a Provider fails; availability with routing enabled cannot be directly compared with availability with routing disabled.

  • BYOK, billing, caching, and data-retention policies differ; the page specifically notes that Anthropic's data-retention rules do not allow zero data retention. Production compliance requires a separate check.

  • Performance data changes with Provider queueing, region, traffic, and the page's time window; this article preserves only the 2026-09-08 snapshot.

Reproduction recommendations

  1. Fix the OpenRouter provider filter, region, routing strategy, model snapshot, prompt length, output limit, and effort.

  2. Send repeated requests to each Provider using the same input set, recording TTFT, end-to-end latency, throughput, structured-output parsing success rate, retries, and fallback.

  3. Report integration-layer metrics separately from capability evaluations; do not use the provider GPQA/TAU snapshots as a substitute for a complete model benchmark.

Curated by Tabbit

This is a third-party source navigator. Model versions, test environments, and personal experience vary; consult the original source.

Claude Fable 5.1

Use and compare models in Tabbit

Claude Fable 5.1

Related reviews

MediaAnthropic News / Introducing Claude Fable 5.1 and Claude Mythos 5.12026-09

Claude Fable 5.1 Official Release: Multiple Benchmarks, Cost Tiers, and Safety Boundaries

MediaArtificial Analysis2026-09-01

Artificial Analysis: Claude Fable 5.1 Intelligence Index, Task Breakdowns, and Cost

MediaSimpleBench

SimpleBench: Claude Fable 5.1's Everyday Reasoning and Human Baseline

CommunityReddit / r/ArtificialInteligence2026-09-02

Reddit community: Skepticism about Fable 5.1 benchmarks and task-tiered experience

Claude Fable 5.1

Related prompts

MediaAnthropic Claude Platform Docs

Claude Fable 5.1 Official Prompting Methods: Writing Density, Batched Tool Calls, and Task Completion

CommunityReddit (r/claude)2026-09-02

Reddit Prompting Techniques: Long-form Analysis, Counterarguments, and Concise Execution

CommunityReddit (r/ClaudeAI)2026-09-05

Reddit case: Fable 5.1 + Blender MCP generates a large world region in one shot