Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiniMax M2.7 · Media / benchmark · Vendor report

MiniMax M2.7 Official Release: SWE-Pro, VIBE-Pro, and Agent Workflow Benchmarks

The MiniMax release page reports 56.22% on SWE-Pro, 76.5 on SWE Multilingual, and 52.7 on Multi-SWE-Bench, alongside a self-feedback workflow.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkVendor reportEdited 2026-09-20

Test conditions

Condition
Model/version: MiniMax M2.7; release page 2026-03.
Condition
Harness/sample: three SWE benchmarks and official workflow; full prompts and repeats unknown.
Condition
Date: page reopened 2026-09-20.

Key data and applicable tasks

One-sentence takeaway

The official results show M2.7 covering a broad range of real engineering, end-to-end project, and Office/Agent tasks: 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2, 1495 on GDPval-AA, and 62.7% on MM Claw, although the complete harness remains undisclosed.

Test environment

  • Model: MiniMax M2.7; official API/Agent and internal Agent harness.

  • Engineering benchmarks: SWE-Pro, VIBE-Pro, Terminal Bench 2, NL2Repo, SWE Multilingual, and Multi SWE Bench.

  • Office/Agent benchmarks: GDPval-AA, Toolathlon, MM Claw (40+ skills/complex work), and MLE Bench Lite (22 ML competitions).

  • Configuration: Each benchmark was run by an official/internal harness; the article does not provide the complete inputs, sampling settings, repeat counts, costs, or confidence intervals for each test.

Inputs/configuration

The official cases include debugging workflows involving production logs/monitoring/deployment timelines/traces/databases/codebases, delivery of Web/Android/iOS/simulated projects, multi-round Word/Excel/PPT editing, research reports, and a TSMC revenue model. Some results came from internal complex skills, Agent Teams, and dynamic tool search.

Results data

  • SWE-Pro: 56.22%; SWE Multilingual: 76.5; Multi SWE Bench: 52.7.

  • VIBE-Pro: 55.6%; Terminal Bench 2: 57.0%; NL2Repo: 39.8%.

  • GDPval-AA: ELO 1495, described by the official source as the highest among open-source models; Toolathlon: 46.3%.

  • MM Claw: 62.7%, described by the official source as close to Sonnet 4.6; across 40+ complex skills, each exceeding 2,000 tokens, skill adherence was 97%.

  • MLE Bench Lite: average medal rate of 66.6% across three 24-hour trials; the best run achieved 9 gold/5 silver/1 bronze.

Conclusion

M2.7 is suitable for Agents that need real-world engineering-system understanding, end-to-end delivery, and tool/skill orchestration. When using the public API, however, treat the official internal harness as an upper-bound reference and remeasure performance with the base model or your own scaffold.

Limitations

  • These are vendor-published benchmarks; the complete prompts, task splits, harnesses, costs, failure samples, and statistical variance are undisclosed.

  • Many figures depend on MiniMax's internal Agent Teams, memory, skills, and dynamic tool search, and cannot be attributed directly to the base model.

  • GDPval-AA, MM Claw, SWE-Pro, and other metrics have different properties and should not be combined into one overall score.

  • “reduced recovery time under three minutes” refers to multiple internal cases reported by the official source, not controlled average performance.

Reproduction steps

  1. Clearly choose the API, open weights, vLLM/SGLang, or a custom Agent harness, and lock the version.

  2. Select coding issues, end-to-end projects, Terminal, office-document, tool-calling, and long-running experiment tasks, and provide the same tools and permissions.

  3. Record prompts, skill/tool definitions, call traces, patches, tests, costs, latency, and human intervention.

  4. Report results separately by benchmark and harness, distinguishing base-model performance from gains due to Agent orchestration.

What this supports

  • Supports citing the three figures as directional results under the official task conditions.

What this does not support

  • Does not support arbitrary-repository fix rates or current API guarantees; task selection and failures are not fully public.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

MiniMax News / Early Echoes of Self-Evolution · MiniMax · Original publication date 2026-03 · Site edit date 2026-09-20

Open original source

MiniMax M2.7

Compare MiniMax M2.7 in Tabbit

Download the Tabbit client to check model access

Related reviews

18 Pieces of Public Evidence for MiniMax M2.7 on BenchLMBenchLM aggregates 18 public pieces of evidence about MiniMax M2.7 and flags different dates, versions, and harnesses; it is an evidence index, not one unified score.Reddit Users' 1,000 Prompts and Coding/Tool-Calling Experience with MiniMax M2.7A Reddit user reports coding and tool-calling experience across about 1,000 prompts, useful for finding parsing failures; it is not a controlled comparison.MiniMax M2.7 Self-Feedback, Memory, and Agent Self-Optimization WorkflowSplit a small fix into plan, implementation, test, and reflection turns; write only verifiable failure causes to memory and compare whether the next run removes the same test failure.MiniMax M2.7 Official Default Prompt and XML Tool-Calling TemplateStarting from MiniMax M2.7’s public default identity prompt, define an XML turn for a code-check tool; parse arguments before execution and verify that the result answers the original request.