MiniMax M3

MiniMax M3 · Reviews and evidence

Which MiniMax M3 conclusions hold up?

Browse public evaluations by topic, source identity, and evidence type. Different versions, tiers, and harnesses are not treated as directly comparable.

This is a third-party source navigator, not a Tabbit test. Use the original source for live metrics; unknown values remain unknown.

Editorial takeaways

Editorial takeaways

A Reddit brownfield Next.js comparison covered API fixes and API additions; M3, MiMo 2.5 Pro, and K2.6 were observed completing tasks, but speed/cost ordering is a single-project observation with undisclosed repeats and harness.

Reddit, r/MiniMaxAI · Read evidence

Reddit users discussed M3 long-horizon coding, context retention, speed, and quotas in Claude Code/OpenCode-style harnesses; task counts, provider snapshots, and unified logs were undisclosed, with a 2026-08-18 collection record.

Reddit, r/MiniMaxAI · Read evidence

Full reviews and related reading

Read the full analysis

Overview · English

MiniMax M3: 1M Context, Coding Power, and the Quota Catch

A source-led MiniMax M3 overview covering M2.7 changes, API and Token Plan access, provider costs, workload fit, Tabbit boundaries, and unknowns.

Selected evidence

CommunityPersonal experience

MiniMax M3: Reddit: Real-Project Benchmark — MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6

A Reddit brownfield Next.js comparison covered API fixes and API additions; M3, MiMo 2.5 Pro, and K2.6 were observed completing tasks, but speed/cost ordering is a single-project observation with undisclosed repeats and harness.

SourceReddit, r/MiniMaxAI
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model and task
MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6 on atomic brownfield Next.js tasks
Client/configuration
Author-run project; full harness, parameters, and repeats undisclosed
Sample
Several API fixes and API additions
Reasoning
CommunityPersonal experience

MiniMax M3: Reddit: MiniMax-M3 Long-Horizon Coding, Speed, and Quota Experience

Reddit users discussed M3 long-horizon coding, context retention, speed, and quotas in Claude Code/OpenCode-style harnesses; task counts, provider snapshots, and unified logs were undisclosed, with a 2026-08-18 collection record.

SourceReddit, r/MiniMaxAI
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Client
Community Claude Code/OpenCode harness; provider and plan affect quotas
Task
Long-horizon coding, context retention, speed, and quota experience
Sample
Personal/commenter reports; task count and unified logs undisclosed
CodingCostSpeed & latency
Media / benchmarkIndependent measurement

MiniMax M3: Google supplement: Artificial Analysis's public metrics for MiniMax-M3

This Google supplement points to public Artificial Analysis metrics for MiniMax-M3; quality, speed, and cost must be read separately within the page version and time window, without inventing provider, tier, sample, or hidden fields.

SourceArtificial Analysis; reached through Google search results
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Platform
Public Artificial Analysis metrics; provider and harness follow the page
Version/tier
Version, reasoning tier, and time window require a fresh check
Metrics
Quality, speed, and cost only as shown; unknown values stay unknown
Capability
Media / benchmarkVendor report

MiniMax M3: Official MiniMax M3 release: coding benchmarks, long context, and real long-task cases

MiniMax’s official material reports coding benchmarks, long context, and long-running agent cases; full prompts, hardware, sample counts, and failures are undisclosed, so it supports vendor positioning rather than independent reproduction or production success rates.

SourceMiniMax official blog
Published2026-06-01
Collected2026-08-18

Unverified: the original source could not be rechecked.

Source
Official MiniMax release material; vendor-reported
Task
Coding benchmarks, long context, and long-running agent cases
Configuration
Full prompts, hardware, sample count, and failures undisclosed
CodingReasoning

All sources

All sources

8 / 8
CommunityPersonal experience

MiniMax M3: Reddit: Real-Project Benchmark — MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6

A Reddit brownfield Next.js comparison covered API fixes and API additions; M3, MiMo 2.5 Pro, and K2.6 were observed completing tasks, but speed/cost ordering is a single-project observation with undisclosed repeats and harness.

SourceReddit, r/MiniMaxAI
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model and task
MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6 on atomic brownfield Next.js tasks
Client/configuration
Author-run project; full harness, parameters, and repeats undisclosed
Sample
Several API fixes and API additions
Reasoning
CommunityPersonal experience

MiniMax M3: Reddit: MiniMax-M3 Long-Horizon Coding, Speed, and Quota Experience

Reddit users discussed M3 long-horizon coding, context retention, speed, and quotas in Claude Code/OpenCode-style harnesses; task counts, provider snapshots, and unified logs were undisclosed, with a 2026-08-18 collection record.

SourceReddit, r/MiniMaxAI
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Client
Community Claude Code/OpenCode harness; provider and plan affect quotas
Task
Long-horizon coding, context retention, speed, and quota experience
Sample
Personal/commenter reports; task count and unified logs undisclosed
CodingCostSpeed & latency
Media / benchmarkIndependent measurement

MiniMax M3: Google supplement: Artificial Analysis's public metrics for MiniMax-M3

This Google supplement points to public Artificial Analysis metrics for MiniMax-M3; quality, speed, and cost must be read separately within the page version and time window, without inventing provider, tier, sample, or hidden fields.

SourceArtificial Analysis; reached through Google search results
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Platform
Public Artificial Analysis metrics; provider and harness follow the page
Version/tier
Version, reasoning tier, and time window require a fresh check
Metrics
Quality, speed, and cost only as shown; unknown values stay unknown
Capability
Media / benchmarkVendor report

MiniMax M3: Official MiniMax M3 release: coding benchmarks, long context, and real long-task cases

MiniMax’s official material reports coding benchmarks, long context, and long-running agent cases; full prompts, hardware, sample counts, and failures are undisclosed, so it supports vendor positioning rather than independent reproduction or production success rates.

SourceMiniMax official blog
Published2026-06-01
Collected2026-08-18

Unverified: the original source could not be rechecked.

Source
Official MiniMax release material; vendor-reported
Task
Coding benchmarks, long context, and long-running agent cases
Configuration
Full prompts, hardware, sample count, and failures undisclosed
CodingReasoning
CommunityPersonal experience

MiniMax M3: Reddit: MiniMax-M3 vs. M2.7 and the Quota Debate

The original author had used M2.7 extensively and considered its quality-to-cost ratio excellent; after trying M3, the main disappointment was the new quota limits rather than the model itself. The comments contain two opposing types of feedback: some users fi。

SourceReddit, r/MiniMaxAI
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model/version
MiniMax-M3; source title “MiniMax M3: Reddit: MiniMax-M3 vs. M2.7 and the Quota Debate”. Exact snapshot follows the original source.
Task/harness
The original author had used M2.7 extensively and considered its quality-to-cost ratio excellent; after trying M3, the main disappointment was the new quota limits rather than the model itself. The comments contain two o The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Cost
CommunityPersonal experience

MiniMax M3: Reddit: Hands-on Measurement of Token Plan Caching and Effective Throughput for MiniMax-M3

This post does not evaluate M3’s intelligence; it measures the Token Plan’s “effective throughput” in an agentic coding scenario. Using OpenCode, OpenRouter BYOK, and cache-hit rates, the author observed that the PAYG caching discount as they understood it did。

SourceReddit, r/MiniMaxAI
PublishedUnknown
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model/version
MiniMax-M3; source title “MiniMax M3: Reddit: Hands-on Measurement of Token Plan Caching and Effective Throughput for MiniMax-M3”. Exact snapshot follows the original source.
Task/harness
This post does not evaluate M3’s intelligence; it measures the Token Plan’s “effective throughput” in an agentic coding scenario. Using OpenCode, OpenRouter BYOK, and cache-hit rates, the author observed that the PAYG ca The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Capability
CommunityPersonal experience

MiniMax M3: X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run

The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost for Fable is modeled. The author's centra。

SourceX
Published2026-08-11
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model/version
MiniMax-M3; source title “MiniMax M3: X: DRACO 100 tasks — four MiniMax-M3 runs plus one synthesis run”. Exact snapshot follows the original source.
Task/harness
The author says that four MiniMax-M3 research runs plus a fifth M3 synthesizer scored 68.1 on the 100-task DRACO benchmark; Fable 5 scored 65.3. The M3 run cost $37 for 100 tasks, while the roughly $250 comparison cost f The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Capability
CommunityPersonal experience

MiniMax M3: X: FutureX real-time forecasting leaderboard — MiniMax-M3-based agent in seventh place

The author reported results for a long-horizon forecasting agent on the FutureX leaderboard: first place for the Kimi K3-based agent, third place for the DeepSeek V4 Pro-based agent, and seventh place for the MiniMax M3-based agent. FutureX was produced by Byt。

SourceX
Published2026-08-04
Collected2026-08-18

Unverified: the original source could not be rechecked.

Model/version
MiniMax-M3; source title “MiniMax M3: X: FutureX real-time forecasting leaderboard — MiniMax-M3-based agent in seventh place”. Exact snapshot follows the original source.
Task/harness
The author reported results for a long-horizon forecasting agent on the FutureX leaderboard: first place for the Kimi K3-based agent, third place for the DeepSeek V4 Pro-based agent, and seventh place for the MiniMax M3- The complete task set, runtime parameters, and review procedure are not fully public.
Sample/date
Source note reviewed 2026-09-20; undisclosed sample count and repeats remain unknown.
Agent

MiniMax M3

Compare MiniMax M3 in Tabbit

Model access, features, and permissions depend on your current client account.