The original post makes only one claim: M3 is inexpensive but highly capable. The comments offer contradictory but more actionable real-world experiences: M3 is cheap and works well as a workhorse for most tasks, but it is slow and may get stuck on complex problems; some users find it more stable than M2.7 during long Agent runs, while others report inconsistent output quality and say that frontier models such as Opus are still needed as a fallback for complex projects.
Consider putting M3 on most implementation, repetitive, and long-context tasks.
For complex debugging, architecture decisions, or tasks where the model gets stuck in a loop, prepare a second model to take over or review.
Speed and reliability are the main weaknesses; “cheap” does not mean a lower total cost for every task.
One commenter said M2.7 drifted after 30–40 turns while M3 retained context more reliably for structured tasks; another said M3 can get stuck on complex problems.
Original post:
Why isn't literally everybody switching from opaque Silicon Valley pricing plans to Chinese ones?
Key comment 1:
M3 runs well and does impressive work for the low cost, but it is slow and can get stuck at complex problems it just… This is a necessary excerpt; read the original source for full context.
The same commenter gave the example that M3 could not correctly handle preserving an Electron tool’s asar packaging and exe unchanged, and also encountered difficulties in an AR game project, while Opus 4.8 quickly found the underlying problems.
Key comment 2:
been using both and M3 feels more stable on long agent runs. M2.7 would drift after 30-40 turns, M3 holds context better… This is a necessary excerpt; read the original source for full context.
Key comment 3:
Now though, MM3 is on par with Kimi 2.6 for intelligence and vision while being much faster. Kimi 2.6's downfall is just… This is a necessary excerpt; read the original source for full context.
This is a community discussion, not a controlled experiment; the post has no standardized task set, run count, or complete logs. The comments’ claims that M3 is “more stable,” “faster,” or “stronger” should all be treated as workflow hypotheses to verify.
MiniMax M3