The author does not trust public vendor benchmarks, so they designed several atomic tasks on a real brownfield project to compare MiniMax-M3, MiMo, and Kimi K2.6. The author says all three completed the tasks, but at different speeds and costs; the post’s TL;DR is that MiMo edges out the others. The result is closer to the experience of everyday development tasks than to a rigorous public benchmark.
The same real project and the same class of tasks are more informative than vendor claims alone.
M3’s advantage does not hold for every one-off coding task; task type, speed, cost, and context retention can change the ranking.
The original author added in the comments that the test used atomic tasks on a Next.js project, such as fixing an API bug or implementing a new API, with the goal of serving their own daily workflow.
A commenter described a different experience: M3 was better than K2.6 at context retention, while K2.6 was faster for single-turn responses; if the model needs to remember more than 20 rounds of tool calls, M3 is more appealing.
I don't trust the new M3 benchmarks, so I made a couple of real tasks on a real, brownfield project,comparing M3 to othe… This is a necessary excerpt; read the original source for full context.
The original author explained further in the comments:
I'm not sure if trust my "benchmark" generally since it's really just some atomic tasks on Next.js. They are just things… This is a necessary excerpt; read the original source for full context.
Another commenter’s hands-on supplement:
ran a similar test on my agent workflow and M3 wins on context retention but K2.6 is faster on single-turn responses. de… This is a necessary excerpt; read the original source for full context.
The post’s charts were published as images, and the current page body does not provide a complete table of per-task scores, times, or costs. This article therefore does not expand “MiMo edges out” into a precise ranking, nor treat the commenter’s personal experience as a general conclusion.
MiniMax M3