Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiniMax M2.7 · Media / benchmark · Independent measurement

18 Pieces of Public Evidence for MiniMax M2.7 on BenchLM

BenchLM aggregates 18 public pieces of evidence about MiniMax M2.7 and flags different dates, versions, and harnesses; it is an evidence index, not one unified score.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkIndependent measurementEdited 2026-09-20

Test conditions

Condition
Model/version: MiniMax M2.7 as labeled by BenchLM; page 2026-08-17.
Condition
Harness/sample: 18 items use mixed settings; not one test set.
Condition
Date: page reopened 2026-09-20.

Key data and applicable tasks

One-sentence takeaway

BenchLM's dynamic evidence ledger lists 18 source-backed results for M2.7, including 56.2% on SWE-Pro, 51.9% on SWE-Rebench, 76.5% on SWE Multilingual, 55.6% on VIBE-Pro, and 46.3% on Toolathlon. It also marks the model as superseded by M3, making it suitable as a historical baseline.

Test environment

  • Directory status: The page marks it as Open Weight, Non-Reasoning, 200K context, Released 2026-03-18, and Superseded.

  • Data date: 2026-08-17; public overall score 63.06/100, ranked #45/218, but this score is BenchLM's custom aggregation.

  • Evidence: The page distinguishes Provider exact, Benchmark exact, and Secondary exact, and provides original links.

  • Coverage: Coding 9, Agentic 6, Knowledge 2, Math 1, and other benchmark categories; this was not one unified experiment.

Inputs/configuration

The page places results from MiniMax's official releases, SWE-Rebench, Vibe Code, React Native, Claw-Eval, Gert Labs, Epoch AI, and others in the same ledger; model versions, tools, sampling, and repeat counts are not fully consistent across rows.

Results data

BenchmarkScorePage evidence
SWE-Rebench51.9%Benchmark exact
SWE-bench Pro56.2%Provider exact, MiniMax official
SWE-bench Verified (mini-swe-agent v2)75.4%Secondary exact
SWE Multilingual76.5%Provider exact
Multi-SWE Bench52.7%Provider exact
VIBE-Pro55.6%Provider exact
Terminal-Bench 2.057.0%Provider exact
Toolathlon46.3%Provider exact
MLE-Bench Lite66.6%Provider exact
MM-ClawBench62.7%Provider exact
NL2Repo39.8%Provider exact

Conclusion

The ledger supports using M2.7 as a strong historical baseline for engineering/tool Agents, but “superseded” means it should be compared with M3 or current models before deployment. Under the same benchmark/harness, the dynamic leaderboard's custom overall score must not be used directly for model selection.

Limitations

  • The BenchLM overall score and ranking are affected by custom weights, coverage, and the dynamic model directory.

  • Provider exact is still vendor-reported; different harnesses under Benchmark exact cannot be combined unconditionally.

  • The page's 75.4% on SWE-bench Verified and MiniMax's prominently released 56.22% on SWE-Pro are different tests and cannot be substituted for one another.

  • The page's current data may change; always save the collection date, source URL, and whether the model is superseded.

Reproduction steps

  1. Open the original benchmark/official source for each row in BenchLM, and fix the version, split, harness, and parameters.

  2. At minimum, rerun SWE-Rebench, SWE-Pro, Terminal-Bench, Toolathlon, and MM Claw-type tasks.

  3. Record resolved/accuracy, tool calls, tokens, latency, cost, and failure samples separately.

  4. Compare M2.7 with M3/other current models using the same harness, and retain M2.7's historical results in the report.

What this supports

  • Supports tracing the 18 items by source date and spotting conclusions newer versions may supersede.

What this does not support

  • Does not support a composite rank, average success rate, or current-capability claim; tasks, versions, prompts, and samples are heterogeneous.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM · BenchLM · Original publication date 2026-08-17 · Site edit date 2026-09-20

Open original source

MiniMax M2.7

Compare MiniMax M2.7 in Tabbit

Download the Tabbit client to check model access

Related reviews

MiniMax M2.7 Official Release: SWE-Pro, VIBE-Pro, and Agent Workflow BenchmarksThe MiniMax release page reports 56.22% on SWE-Pro, 76.5 on SWE Multilingual, and 52.7 on Multi-SWE-Bench, alongside a self-feedback workflow.Reddit Users' 1,000 Prompts and Coding/Tool-Calling Experience with MiniMax M2.7A Reddit user reports coding and tool-calling experience across about 1,000 prompts, useful for finding parsing failures; it is not a controlled comparison.MiniMax M2.7 Self-Feedback, Memory, and Agent Self-Optimization WorkflowSplit a small fix into plan, implementation, test, and reflection turns; write only verifiable failure causes to memory and compare whether the next run removes the same test failure.MiniMax M2.7 Official Default Prompt and XML Tool-Calling TemplateStarting from MiniMax M2.7’s public default identity prompt, define an XML turn for a code-check tool; parse arguments before execution and verify that the result answers the original request.