MiniMax M3 review navigator
Official benchmarks, independent analysis, and community reports about MiniMax M3, clearly separated from Tabbit's own testing.
Media
2 source-checked resourcesGoogle supplement: Artificial Analysis's public metrics for MiniMax-M3
Google results show that Artificial Analysis's MiniMax-M3 page compares model quality, price, output speed, and latency. A Google snippet gives an Intelligence Index of about 45 and a Coding Index of about 58.6; different result cards also showed 55 as an olde。
Official MiniMax M3 release: coding benchmarks, long context, and real long-task cases
Model: MiniMax M3, a MoE model with approximately 428B total parameters and approximately 23B active parameters; MiniMax Sparse Attention (MSA) supports up to a 1M context; native image/video input and computer use.。
Community
6 source-checked resourcesReddit: Real-Project Benchmark — MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6
The author does not trust public vendor benchmarks, so they designed several atomic tasks on a real brownfield project to compare MiniMax-M3, MiMo, and Kimi K2.6. The author says all three completed the tasks, but at different speeds and costs; the post’s TL;D。
Reddit: MiniMax-M3 Long-Horizon Coding, Speed, and Quota Experience
The original post makes only one claim: M3 is inexpensive but highly capable. The comments offer contradictory but more actionable real-world experiences: M3 is cheap and works well as a workhorse for most tasks, but it is slow and may get stuck on complex pro。
Reddit: MiniMax-M3 vs. M2.7 and the Quota Debate
The original author had used M2.7 extensively and considered its quality-to-cost ratio excellent; after trying M3, the main disappointment was the new quota limits rather than the model itself. The comments contain two opposing types of feedback: some users fi。
Reddit: Hands-on Measurement of Token Plan Caching and Effective Throughput for MiniMax-M3
This post does not evaluate M3’s intelligence; it measures the Token Plan’s “effective throughput” in an agentic coding scenario. Using OpenCode, OpenRouter BYOK, and cache-hit rates, the author observed that the PAYG caching discount as they understood it did。
X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run
The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's centra。
X: FutureX real-time forecasting leaderboard — MiniMax-M3-based agent in seventh place
The author reported results for a long-horizon forecasting agent on the FutureX leaderboard: first place for the Kimi K3-based agent, third place for the DeepSeek V4 Pro-based agent, and seventh place for the MiniMax M3-based agent. FutureX was produced by Byt。
MiniMax M3
Use and compare models in Tabbit
Official benchmarks, independent analysis, and community reports about MiniMax M3, clearly separated from Tabbit's own testing.