MiniMax M3 · Community source · Personal experience
The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's centra。
Unverified: the original source could not be rechecked. Historical figures below are not current verified results.
The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's central point is that multiple cheap copies can provide useful diversity.
In this experimental setup, combining multiple inexpensive runs with a synthesizer outscored a single Fable 5.
The 68.1 score cannot be attributed to a single M3 run; the result came from the system of “four research runs + a fifth synthesizer.”
The cost advantage is approximately 37/250 = 14.8%, but the Fable cost is modeled rather than a like-for-like billed amount.
This result supports evaluating “multiple sampling/model orchestration,” rather than only a single response from a single model.
Four cheap runs beat one frontier model. Four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 10… This is a necessary excerpt; read the original source for full context.
Title of a related link by the same author:
Four copies of a cheap model beat Fable at 1/7 the price
The original post did not disclose the complete task set, scoring details, outputs from each run, or Fable's actual bill. It should be treated as an experimental lead about “whether multiple runs are worthwhile,” not as a complete reproducible benchmark report.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X · @jperla (Joseph Perla) · Original publication date 2026-08-11 · Site edit date 2026-09-20
Open original sourceMiniMax M3
Download the Tabbit client to check model access