Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiniMax M3 · Media / benchmark · Vendor report

MiniMax M3: Official MiniMax M3 release: coding benchmarks, long context, and real long-task cases

MiniMax’s official material reports coding benchmarks, long context, and long-running agent cases; full prompts, hardware, sample counts, and failures are undisclosed, so it supports vendor positioning rather than independent reproduction or production success rates.

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Media / benchmarkVendor reportEdited 2026-09-20

Test conditions

Source
Official MiniMax release material; vendor-reported
Task
Coding benchmarks, long context, and long-running agent cases
Configuration
Full prompts, hardware, sample count, and failures undisclosed
Date
Release and collection dates are separate; current status requires a fresh check

Key data and applicable tasks

Test environment

  • Model: MiniMax M3, a MoE model with approximately 428B total parameters and approximately 23B active parameters; MiniMax Sparse Attention (MSA) supports up to a 1M context; native image/video input and computer use.

  • Public benchmarks: SWE-Bench Pro, Terminal-Bench 2.1, SWE-fficiency, KernelBench Hard, MCP Atlas, and PostTrainBench.

  • Long-task cases: reproducing a paper, optimizing NVIDIA Hopper FP8 GEMM, and having the model autonomously post-train four foundational models.

  • Evaluation method: the official internal infrastructure, with some tasks using harnesses such as Claude Code/Terminus 2; the official description says SWE-Bench Verified was run four times and averaged, while other metrics used the corresponding sandbox and timeout settings.

Inputs/configuration

  • The official benchmarks used M3's coding/Agent configuration; Terminal-Bench 2.1 used an 8C16G sandbox, a two-hour timeout, a 128K maximum output, and Terminus 2 scaffolding.

  • Paper reproduction: the paper, code, and experiment logs were placed in the long-task context, and the model ran autonomously for nearly 12 hours.

  • CUDA optimization: the task description, benchmark script, and a Triton skeleton that could not be run directly were provided, without a reference high-performance implementation; the model iterated using benchmark feedback.

Result data

Benchmark/caseOfficial result
SWE-Bench Pro59.0%
Terminal-Bench 2.166.0%
SWE-fficiency34.8%
KernelBench Hard28.8%
MCP Atlas74.2%
PostTrainBench0.37 (Opus 4.7: 0.42; GPT-5.5: 0.39)
  • Paper reproduction: nearly 12 hours, 18 commits, and 23 experiment figures; the core experiments were completed and the main trends reproduced.

  • CUDA FP8 GEMM: approximately 24 hours, 147 benchmark submissions, and 1,959 tool calls; peak utilization rose from 7.6% to 71.3%, or 9.4× relative to the initial version.

  • The official source says that under a 1M context, MSA requires approximately 1/20 of the per-token computation of the previous generation, with more than 9× prefill acceleration and more than 15× decode acceleration; these are architecture/service-side figures, not a general end-user throughput commitment.

Conclusion

The evidence for M3 is concentrated in long-horizon coding, iterative tool feedback, and multimodal Agents, rather than simple single-turn chat. The official cases show that it can continue exploring for an extended period under clear benchmark feedback; its benchmark scores are in a usable frontier range, but they cannot be interpreted independently of the harness and internal evaluation methodology.

Limitations

  • All results are vendor-reported; some comparison scores come from official leaderboards or different harnesses. The official source did not disclose all raw trajectories, failure samples, or random seeds.

  • Paper reproduction and CUDA optimization are selected cases and cannot represent the success rate on ordinary repositories.

  • The 1M context and MSA acceleration figures do not mean that every API provider offers the same context, latency, or pricing.

Reproduction steps

  1. For public coding tasks, fix the sandbox, tools, timeout, output limit, and harness, and run at least three times.

  2. Reproduce experimental long tasks separately: record every submission, benchmark score, tool call, token count, wall-clock time, and human intervention.

  3. Report model capability, harness orchestration, and hardware/service throughput separately; do not substitute selected cases for task-set statistics.

  4. For tasks requiring image/video input, verify that the actual API endpoint provides native multimodality and the same billing terms.

Original evidence and data

  • The official release page clearly lists the five coding/Agent benchmarks above and the PostTrainBench comparison values.

  • The official release page discloses the timing, submission counts, tool-call counts, and result changes for the 12-hour paper reproduction and 24-hour CUDA optimization.

  • The official evaluation-method description explains the main harness/sandbox settings for SWE-Bench Verified, Terminal-Bench 2.1, and NL2Repo.

Source excerpt or observation (compliance short quote only)

  • The official positioning puts “frontier coding, 1M context, native multimodality” in the same model; actual selection still needs to return to the specific toolchain and feedback loop.

What this supports

  • Supports using official benchmarks and long-task cases for product positioning.

What this does not support

  • Does not support independent reproduction, general success rates, or Tabbit end-to-end availability.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

MiniMax official blog · MiniMax · Original publication date 2026-06-01 · Site edit date 2026-09-20

Open original source

MiniMax M3

Compare MiniMax M3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

MiniMax M3: 1M Context, Coding Power, and the Quota Catch

A source-led MiniMax M3 overview covering M2.7 changes, API and Token Plan access, provider costs, workload fit, Tabbit boundaries, and unknowns.

Related reviews

MiniMax M3: Reddit: Real-Project Benchmark — MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6A Reddit brownfield Next.js comparison covered API fixes and API additions; M3, MiMo 2.5 Pro, and K2.6 were observed completing tasks, but speed/cost ordering is a single-project observation with undisclosed repeats and harness.MiniMax M3: Reddit: MiniMax-M3 Long-Horizon Coding, Speed, and Quota ExperienceReddit users discussed M3 long-horizon coding, context retention, speed, and quotas in Claude Code/OpenCode-style harnesses; task counts, provider snapshots, and unified logs were undisclosed, with a 2026-08-18 collection record.MiniMax M3: Google supplement: Artificial Analysis's public metrics for MiniMax-M3This Google supplement points to public Artificial Analysis metrics for MiniMax-M3; quality, speed, and cost must be read separately within the page version and time window, without inventing provider, tier, sample, or hidden fields.MiniMax M3: Reddit: MiniMax-M3 vs. M2.7 and the Quota DebateThe original author had used M2.7 extensively and considered its quality-to-cost ratio excellent; after trying M3, the main disappointment was the new quota limits rather than the model itself. The comments contain two opposing types of feedback: some users fi。MiniMax M3: Google Supplement: Integration Prompting for Official MiniMax M3 with Claude Code / OpenCodeTurn Google Supplement: Integration Prompting for Official MiniMax M3 with Claude Code / OpenCode into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.MiniMax M3: Reddit: MiniMax-M3 Routing and Orchestration for Long Tasks in Claude CodeTurn Reddit: MiniMax-M3 Routing and Orchestration for Long Tasks in Claude Code into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.MiniMax M3: Reddit: Caching, Context, and Billing Verification in an M3 Agent Prompt WorkflowTurn Reddit: Caching, Context, and Billing Verification in an M3 Agent Prompt Workflow into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.MiniMax M3: MiniMax Official: M-Series Prompting Best PracticesTurn MiniMax Official: M-Series Prompting Best Practices into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.