The original author had used M2.7 extensively and considered its quality-to-cost ratio excellent; after trying M3, the main disappointment was the new quota limits rather than the model itself. The comments contain two opposing types of feedback: some users find M3 smarter and more stable for long-running Agents, while others find it slow, inconsistent in output quality, and too quick to consume quotas. The post is useful as a counterexample showing that “model capability” and “effective throughput under a plan” need to be evaluated separately.
Evaluate M3 on two levels: model quality, and the usable workload enabled by the plan, caching, and quota.
Some commenters recommended disabling thinking to reduce consumption, or using M2.7 as the execution model and M3 as the planning model.
M3’s context retention may help with complex, structured tasks; quality feedback is less consistent for creative work.
The same model can feel substantially different under different harnesses, thinking settings, and caching implementations.
I've been a heavy user of Minimax M2.7 over the past few months and honestly thought it was one of the most underrated m… This is a necessary excerpt; read the original source for full context.
Key comment:
You can disable thinking in M3, that'll get the burn rate closer to M2.7. Or just use M2.7 for work and M3 for planning… This is a necessary excerpt; read the original source for full context.
Another account with the opposite experience:
been using both and M3 feels more stable on long agent runs. M2.7 would drift after 30-40 turns, M3 holds context better… This is a necessary excerpt; read the original source for full context.
The original post provides no standardized task set or billing logs; conclusions about quotas, caching, and quality are clearly disputed and cannot be treated directly as product-pricing facts.
MiniMax M3