The author reported results for a long-horizon forecasting agent on the FutureX leaderboard: first place for the Kimi K3-based agent, third place for the DeepSeek V4 Pro-based agent, and seventh place for the MiniMax M3-based agent. FutureX was produced by ByteDance Seed together with Stanford, Princeton, and Fudan; its questions concern real events that have not yet occurred. Predictions are submitted first and scored against the actual outcomes after the events resolve. The author emphasized that the same framework was used, with only the three models serving as the “brain” being swapped.
M3 entered the top ten in this specific forecasting agent and harness, indicating that it can participate in long-horizon retrieval, evidence synthesis, and probabilistic forecasting.
This is a system result from “model + harness + retrieval/adjudication mechanism,” not a bare-model score.
In a reply, the original author explicitly said that the leaderboard score measures the capability of the entire harness system; comparing the models themselves requires controlling the harness variable and conducting a stable, reproducible comparative experiment.
For the past two months, we've been quietly working on one thing: teaching AI to predict the future. Today we can finall… This is a necessary excerpt; read the original source for full context.
Author's reply:
I think the leaderboard score is not a measurement of model capability, but of the capability of the entire harness syst… This is a necessary excerpt; read the original source for full context.
This is an Agent result from the author's team. The standalone MiniMax-M3 score, complete task set, and comparison harness were not disclosed. It cannot support the claim that M3 ranks seventh across all forecasting tasks.
MiniMax M3