The public SimpleBench leaderboard shows Fable 5.1 scoring 86.6%, above the 83.7% average baseline of the project's nine human participants. This indicates strong performance on spatial/temporal reasoning, social common sense, and language-trap questions, but the human-baseline sample is very small and does not prove general capability.
Suitable tasks: Multiple-choice reasoning that requires recognizing everyday situations, spatial/temporal relationships, social intelligence, and language traps.
Unsuitable tasks: Selecting a model for coding, tool use, long-horizon Agents, scientific research, or professional knowledge work; SimpleBench is a pure text multiple-choice benchmark and cannot replace evals for those tasks.
Applicable model version: The leaderboard lists Claude Fable 5.1, with the score labeled AVG@5; the page does not publicly show the specific effort setting in the leaderboard row.
Applicable client, Agent, or API: The page provides a public dataset, code, and an online test; details of the model evaluation run should be based on the project's code and technical report.
Recommended reasoning tier and parameters: The source page does not disclose Fable 5.1's effort, temperature, or complete API parameters; when reproducing the evaluation, do not assume max or the default tier.
SimpleBench is a text multiple-choice benchmark with more than 200 questions, covering spatiotemporal reasoning, social intelligence, and “language-adversarial robustness” (trick questions). Each question has 6 options.
The page states that nine non-expert human participants make up an average baseline of 83.7%; each participant completed 25 randomly selected questions, covering all 200+ questions. Selection criteria were English native speakers with at least a high-school level of mathematics; IQ and general reasoning ability were not controlled for.
The current leaderboard uses Score (AVG@5). Fable 5.1 scores 86.6%, GPT-6 Astra Pro 86.5%, the human baseline 83.7%, GPT-6 Astra 83.6%, Gemini 3.8 Flash 82.4%, and Fable 5 81.9%.
The project publishes the simple_bench_public.json dataset and a GitHub code link; the page also provides a technical report, offering entry points for further review.
Fable 5.1 exceeds the human average baseline by approximately 2.9 percentage points on this set of everyday reasoning questions that emphasize situations that “seem simple but can easily mislead people through language or common-sense traps.” It is suitable for inclusion in regression suites for everyday judgment, semantic traps, and social common sense, but this result cannot be extrapolated to mean that it is “more intelligent than humans overall.”
The human baseline has only 9 participants, and the project itself explicitly describes the sample as small; participant selection is subject to bias, so 83.7% does not equal the population average.
The leaderboard's AVG@5 indicates averaging across multiple runs, but the current page does not publicly show the prompt, temperature, effort, model snapshot, or random seed for each sample in the leaderboard row; reproduction requires consulting the project's code and technical report and recording the versions.
The questions are text multiple-choice questions and do not measure tool use, code execution, knowledge retrieval, output quality, or real-world project success rates.
The leaderboard will change as new models are added; this article records the snapshot collected on 2026-09-08 and should not be treated as a permanent ranking.
“Exceeds the human baseline” applies only to the average-score comparison on this benchmark and cannot be directly converted into a conclusion about general AGI or real-world work efficiency.
Download the public question set and code from the project page, and lock the commit, question-set version, and AVG@5 evaluation method.
Specify the Fable 5.1 model snapshot, effort, temperature, max tokens, and whether multiple samples are allowed; save the five outputs for each question and the final aggregate score.
Evaluate models such as Fable 5, Opus 5, and GPT-5.6 Sol using exactly the same question set and parameters, and report each subscore and confidence interval.
Use the results only to assess “everyday reasoning/trick-question” capability, and present them separately from Agent, coding, and knowledge-work benchmarks.
Claude Fable 5.1