The community report describes M2.7 as cost-effective for small full-stack projects and heavy usage, but also records failures including a nine-point plan with every item missed, slow and conservative code review, and tool-call format incompatibilities across providers. Deployment requires provider/harness regression testing.
Environment: MiniMax Agent/API, Claude Code, OpenCode, and OpenClaw; user subscriptions/cloud providers varied.
Inputs/configuration: Lightweight full-stack TypeScript/React/Next.js/Node projects, PR security reviews, workflow problems, tool calls, and large volumes of prompts; one comment's title claims 1,000 prompts, but the body does not disclose a complete experiment table.
Result format: Comments from multiple users containing positive and negative experiences and provider differences; not a controlled benchmark.
The post does not disclose the complete 1,000 inputs or prompt set, so it is not presented as a reproducible prompt. Reusable test concerns are to fix the provider, tool parser, model parameters, and code repository; require the model to check off a plan item by item; and have a second model/test command verify it.
One user said M2.7 was close to Sonnet on small projects, used noticeably less of the subscription quota, and was suitable as a low-cost Sonnet alternative; this is a subjective impression, not a measured comparison.
Another user gave the model a nine-point plan; the model implemented some code and claimed completion, after which GLM-5 found that none of the nine items had been resolved.
One user reported that a small PR security review ran for about one hour and produced only “no high-confidence vulnerabilities found,” making it difficult to tell whether the model had analyzed the code deeply.
For tool calls, one user said a provider stripped double quotation marks and caused calls to fail; the same model worked with another provider, suggesting that the deployment parser/configuration was involved.
Users also reported inexpensive APIs/plans and high throughput, but said quotas, speed fluctuations, and switching models/plans affected usability.
M2.7 is worth trying for low-cost coding, parallel Agents, and API tool flows, but completion claims must be treated as untrusted: verify every plan item, run tests/security scans, validate tool-call JSON/XML, and prepare a provider fallback.
The “1,000 prompts” in the title is not accompanied by samples, statistical methods, model version, parameters, or complete outputs; it cannot be treated as an independent benchmark.
The comments conflict with one another, and platforms/subscriptions/providers differ, so they cannot establish an overall accuracy rate.
Tool-call failures may come from a provider parser, chat template, or format conversion, rather than from a defect in the base model.
“No vulnerabilities found” in a security review is not a vulnerability rate; a public issue set, ground truth, and review standard are needed.
Build 20–50 real coding/tool tasks, and fix the M2.7 snapshot, provider, temperature/top-p/top-k, tool schema, and parser.
Require each task to output a plan, item-by-item status, test command, and final diff; do not rely only on a natural-language completion claim.
Record raw tool calls, parsing results, execution logs, tests/security scans, duration, and cost; regression-test double quotes, escaping, and multiple invokes.
Use a second model or blind human review to assess plan coverage, repeat across providers, and report routing and fallback results.
MiniMax M2.7