OpenRouter's live page shows Fable 5.1 delivering roughly 44–56 tok/s in one-week average throughput across multiple Providers, with roughly 4.18–6.01 seconds of average first-chunk/request latency, and provides routing snapshots for GPQA Diamond, TAU-Bench, and structured-output error rates. It is suitable for choosing an integration layer, not for replacing controlled capability evaluations.
Suitable tasks: Comparing latency, throughput, availability, and structured-output performance when connecting to the same model through Azure, Anthropic, Amazon Bedrock BYOK, and Google Vertex.
Unsuitable tasks: Inferring the model's reasoning ability, long-horizon Agent success rates, or the end-to-end experience across different clients; Provider routing and traffic composition affect the results.
Applicable model version: anthropic/claude-fable-5.1 on OpenRouter.
Applicable clients, Agents, or APIs: OpenRouter routing; the page lists high-traffic applications such as Hermes Agent, Claude Code, and Kilo Code, but this is not a controlled comparison experiment for those applications.
Recommended reasoning tier and parameters: The page does not disclose the effort, prompt, output length, or sampling settings corresponding to this telemetry; production integrations should fix these variables independently.
The page displays P50 runtime data for each Provider:
Standard endpoint pricing is $10/million tokens for input, $50/million tokens for output, and $0.25/million tokens for cache reads.
Provider P50: Azure latency 14.96s and throughput 56 tok/s; Anthropic latency 5.97s, throughput 45 tok/s, and availability 99.91%; Amazon Bedrock (BYOK) latency 11.58s and throughput 53 tok/s; Google Vertex latency 4.84s, throughput 65 tok/s, and availability 99.35%.
Average throughput over the past week: Vertex 56 tok/s, Bedrock 46 tok/s, and Anthropic 44 tok/s; average first-chunk/request latency: Vertex 4.18s, Anthropic 5.96s, and Azure 6.01s; end-to-end average latency was 12.51s, 16.20s, and 16.15s, respectively.
AutoExacto Benchmarks: GPQA Diamond—automatic routing 90.9%, Anthropic 86.5%, Vertex 84.9%, and Azure 86.3%; TAU-Bench—Anthropic 79.3%, Vertex 76.7%, and Azure 74.7%. The page does not provide the complete question sets or repetition counts for the automatic-routing and Provider results.
Structured-output error rates: Anthropic averaged 8.33%, Bedrock 33.12%, and Azure 38.51%; cache hit rates were approximately 85.26%–86.82%.
At collection time, the page showed OpenRouter uptime of 100% and availability of 99.92% over the past 3 days, while "no-routing" availability, which does not bypass Provider failures, was 98.25%.
OpenRouter's evidence is primarily useful for deployment-layer decisions: Vertex led in throughput and latency in this page snapshot, while the Anthropic endpoint had a lower structured-output error rate; if a task depends on JSON/structured calls, Provider differences may affect stability more than the model name itself. In production, fix the provider filter, routing strategy, and retry method before measuring again.
This is platform live telemetry, not a public controlled benchmark report; the page does not provide the complete input set, output-length distribution, effort, sampling parameters, or statistical confidence intervals.
P50, past-week averages, and past-three-days uptime use different time windows and cannot be treated as one unified metric.
OpenRouter automatically switches when a Provider fails; availability with routing enabled cannot be directly compared with availability with routing disabled.
BYOK, billing, caching, and data-retention policies differ; the page specifically notes that Anthropic's data-retention rules do not allow zero data retention. Production compliance requires a separate check.
Performance data changes with Provider queueing, region, traffic, and the page's time window; this article preserves only the 2026-09-08 snapshot.
Fix the OpenRouter provider filter, region, routing strategy, model snapshot, prompt length, output limit, and effort.
Send repeated requests to each Provider using the same input set, recording TTFT, end-to-end latency, throughput, structured-output parsing success rate, retries, and fallback.
Report integration-layer metrics separately from capability evaluations; do not use the provider GPQA/TAU snapshots as a substitute for a complete model benchmark.
Claude Fable 5.1