Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiniMax M3 · Community source · Personal experience

MiniMax M3: X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run

The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's centra。

Unverified: the original source could not be rechecked. Historical figures below are not current verified results.

Community sourcePersonal experienceEdited 2026-09-20

Test conditions

Model/version
MiniMax-M3; source title “MiniMax M3: X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run”. Exact snapshot follows the original source.
Task/harness
The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost f The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.

Key data and applicable tasks

Summary

The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's central point is that multiple cheap copies can provide useful diversity.

Available conclusions

  • In this experimental setup, combining multiple inexpensive runs with a synthesizer outscored a single Fable 5.

  • The 68.1 score cannot be attributed to a single M3 run; the result came from the system of “four research runs + a fifth synthesizer.”

  • The cost advantage is approximately 37/250 = 14.8%, but the Fable cost is modeled rather than a like-for-like billed amount.

  • This result supports evaluating “multiple sampling/model orchestration,” rather than only a single response from a single model.

Original post

Four cheap runs beat one frontier model. Four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 10… This is a necessary excerpt; read the original source for full context.

Title of a related link by the same author:

Four copies of a cheap model beat Fable at 1/7 the price

Limitations

The original post did not disclose the complete task set, scoring details, outputs from each run, or Fable's actual bill. It should be treated as an experimental lead about “whether multiple runs are worthwhile,” not as a complete reproducible benchmark report.

What this supports

  • Supports the source-specific observation in “MiniMax M3: X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run”: The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the

What this does not support

  • Does not support a general capability, production success-rate, or current-ranking claim from “MiniMax M3: X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run”; the source lacks a controlled task set, provider snapshot, and repeated independent retest.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

X · @jperla (Joseph Perla) · Original publication date 2026-08-11 · Site edit date 2026-09-20

Open original source

MiniMax M3

Compare MiniMax M3 in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Overview · English

MiniMax M3: 1M Context, Coding Power, and the Quota Catch

A source-led MiniMax M3 overview covering M2.7 changes, API and Token Plan access, provider costs, workload fit, Tabbit boundaries, and unknowns.

Related reviews

MiniMax M3: Google supplement: Artificial Analysis's public metrics for MiniMax-M3This Google supplement points to public Artificial Analysis metrics for MiniMax-M3; quality, speed, and cost must be read separately within the page version and time window, without inventing provider, tier, sample, or hidden fields.MiniMax M3: Reddit: Hands-on Measurement of Token Plan Caching and Effective Throughput for MiniMax-M3This post does not evaluate M3’s intelligence; it measures the Token Plan’s “effective throughput” in an agentic coding scenario. Using OpenCode, OpenRouter BYOK, and cache-hit rates, the author observed that the PAYG caching discount as they understood it did。MiniMax M3: Reddit: Real-Project Benchmark — MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6A Reddit brownfield Next.js comparison covered API fixes and API additions; M3, MiMo 2.5 Pro, and K2.6 were observed completing tasks, but speed/cost ordering is a single-project observation with undisclosed repeats and harness.MiniMax M3: Reddit: MiniMax-M3 Long-Horizon Coding, Speed, and Quota ExperienceReddit users discussed M3 long-horizon coding, context retention, speed, and quotas in Claude Code/OpenCode-style harnesses; task counts, provider snapshots, and unified logs were undisclosed, with a 2026-08-18 collection record.MiniMax M3: Reddit: MiniMax-M3 Routing and Orchestration for Long Tasks in Claude CodeTurn Reddit: MiniMax-M3 Routing and Orchestration for Long Tasks in Claude Code into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.MiniMax M3: MiniMax Official: M-Series Prompting Best PracticesTurn MiniMax Official: M-Series Prompting Best Practices into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.MiniMax M3: Google Supplement: Integration Prompting for Official MiniMax M3 with Claude Code / OpenCodeTurn Google Supplement: Integration Prompting for Official MiniMax M3 with Claude Code / OpenCode into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.MiniMax M3: MiniMax Official M3 Long-Running Agent Workflow: Paper Reproduction and Producer/Verifier Self-CheckingTurn MiniMax Official M3 Long-Running Agent Workflow: Paper Reproduction and Producer/Verifier Self-Checking into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.