BenchLM's dynamic evidence ledger lists 18 source-backed results for M2.7, including 56.2% on SWE-Pro, 51.9% on SWE-Rebench, 76.5% on SWE Multilingual, 55.6% on VIBE-Pro, and 46.3% on Toolathlon. It also marks the model as superseded by M3, making it suitable as a historical baseline.
Directory status: The page marks it as Open Weight, Non-Reasoning, 200K context, Released 2026-03-18, and Superseded.
Data date: 2026-08-17; public overall score 63.06/100, ranked #45/218, but this score is BenchLM's custom aggregation.
Evidence: The page distinguishes Provider exact, Benchmark exact, and Secondary exact, and provides original links.
Coverage: Coding 9, Agentic 6, Knowledge 2, Math 1, and other benchmark categories; this was not one unified experiment.
The page places results from MiniMax's official releases, SWE-Rebench, Vibe Code, React Native, Claw-Eval, Gert Labs, Epoch AI, and others in the same ledger; model versions, tools, sampling, and repeat counts are not fully consistent across rows.
| Benchmark | Score | Page evidence |
|---|---|---|
| SWE-Rebench | 51.9% | Benchmark exact |
| SWE-bench Pro | 56.2% | Provider exact, MiniMax official |
| SWE-bench Verified (mini-swe-agent v2) | 75.4% | Secondary exact |
| SWE Multilingual | 76.5% | Provider exact |
| Multi-SWE Bench | 52.7% | Provider exact |
| VIBE-Pro | 55.6% | Provider exact |
| Terminal-Bench 2.0 | 57.0% | Provider exact |
| Toolathlon | 46.3% | Provider exact |
| MLE-Bench Lite | 66.6% | Provider exact |
| MM-ClawBench | 62.7% | Provider exact |
| NL2Repo | 39.8% | Provider exact |
The ledger supports using M2.7 as a strong historical baseline for engineering/tool Agents, but “superseded” means it should be compared with M3 or current models before deployment. Under the same benchmark/harness, the dynamic leaderboard's custom overall score must not be used directly for model selection.
The BenchLM overall score and ranking are affected by custom weights, coverage, and the dynamic model directory.
Provider exact is still vendor-reported; different harnesses under Benchmark exact cannot be combined unconditionally.
The page's 75.4% on SWE-bench Verified and MiniMax's prominently released 56.22% on SWE-Pro are different tests and cannot be substituted for one another.
The page's current data may change; always save the collection date, source URL, and whether the model is superseded.
Open the original benchmark/official source for each row in BenchLM, and fix the version, split, harness, and parameters.
At minimum, rerun SWE-Rebench, SWE-Pro, Terminal-Bench, Toolathlon, and MM Claw-type tasks.
Record resolved/accuracy, tool calls, tokens, latency, cost, and failure samples separately.
Compare M2.7 with M3/other current models using the same harness, and retain M2.7's historical results in the report.
MiniMax M2.7