The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's central point is that multiple cheap copies can provide useful diversity.
In this experimental setup, combining multiple inexpensive runs with a synthesizer outscored a single Fable 5.
The 68.1 score cannot be attributed to a single M3 run; the result came from the system of “four research runs + a fifth synthesizer.”
The cost advantage is approximately 37/250 = 14.8%, but the Fable cost is modeled rather than a like-for-like billed amount.
This result supports evaluating “multiple sampling/model orchestration,” rather than only a single response from a single model.
Four cheap runs beat one frontier model. Four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 10… This is a necessary excerpt; read the original source for full context.
Title of a related link by the same author:
Four copies of a cheap model beat Fable at 1/7 the price
The original post did not disclose the complete task set, scoring details, outputs from each run, or Fable's actual bill. It should be treated as an experimental lead about “whether multiple runs are worthwhile,” not as a complete reproducible benchmark report.
MiniMax M3