Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Prompts and workflows

MiniMax M3 · workflow

MiniMax M3: MiniMax Official M3 Long-Running Agent Workflow: Paper Reproduction and Producer/Verifier Self-Checking

Turn MiniMax Official M3 Long-Running Agent Workflow: Paper Reproduction and Producer/Verifier Self-Checking into an executable task with explicit inputs, environment, and boundaries; see the detail page for steps and limits.

Source not verifiedMiniMax API, MiniMax Code, or a compatible agent harness

Prerequisites and inputs

  • Task goal
  • Source material
  • Output format
  • Acceptance criteria

Complete templates

Editorial adaptation: Producer/Verifier reproduction

Tabbit editorial adaptation; not the original source prompt
Goal:
{{GOAL}}
Materials and commands:
{{MATERIALS}}
Independent checks:
{{SUCCESS_CHECKS}}
Budget and handoff limit:
{{BUDGET}}
Producer runs one change; Verifier checks raw results; Coordinator records continue, rollback, replan, or stop.

Replace before running: {{GOAL}}, {{MATERIALS}}, {{SUCCESS_CHECKS}}, {{BUDGET}}

Prerequisites

Prepare papers, code, data, logs, runnable commands, hardware limits, a cost ceiling, and separate Producer/Verifier/Coordinator permissions.

Steps

  1. Define an artifact, command, success metric, stop condition, and handoff point per phase.

  2. Producer makes one minimal change and runs the benchmark; Verifier independently checks code, data, charts, and conclusions.

  3. Coordinator continues, rolls back, replans, or stops only on new evidence, saving commits, logs, and unreproduced claims.

Checks and fixes

Do not repeat without new feedback; narrow or roll back after failed verification, and hand off when budget, permission, or safety limits are exceeded.

Source boundary

Official cases describe long-running experiments and roles but do not disclose the full system prompt, hardware, failures, or production success rate.

Read the source research notes

One-sentence takeaway

Long M3 tasks should give the Agent the paper, code, logs, and executable verification together, then let it advance through a “plan—execute—feedback—replan” loop; MiniMax Code's Producer + Verifier structure can serve as a template for multi-Agent self-checking.

Use cases

  • Suitable tasks: Paper reproduction, CUDA/kernel optimization, cross-file engineering iteration, long-running experiments, and Agent tasks requiring continuous tool feedback.

  • Unsuitable tasks: Open-ended creation without executable tests or feedback signals, and simple tasks requiring low-latency instant answers.

  • Applicable model versions: MiniMax M3; on the product side, MiniMax Code + M3.

  • Applicable clients, Agents, or APIs: MiniMax Code, Agent harnesses based on OpenCode/Pi, and self-built workflows with runnable commands and tests.

  • Recommended reasoning tier and parameters: Enable thinking for complex tasks; disable thinking for short conversations or code completion in exchange for speed. Use official defaults for other parameters.

Ready-to-use content

Task: [Clearly state the research/engineering goal and success metrics]
Inputs: [Paper, code, data, logs, hardware, and constraints]
Plan: Break the work into verifiable stages first, specifying each stage's outputs, commands, and stopping conditions.
Execution: Run the minimum verification after each change; record the result, reason for failure, and next hypothesis.
Iteration: Reorder the plan based on test/benchmark feedback; do not repeat the same attempt without new evidence.
Verification: Independently review the implementation, results, and charts; list unreproduced conclusions and remaining risks.
Deliverable: Submit the changes, experiment results, reproducible commands, charts, and boundaries of the conclusions.

Multi-Agent version:

Producer: Propose the implementation and run experiments, outputting changes and evidence.
Verifier: Independently check requirements, tests, data, and results, and identify failures or overfitting.
Coordinator: Decide whether to continue, roll back, revise the plan, or finish based on Verifier feedback.

Test/workflow steps

  1. Put the task materials in an accessible directory, then have the model list the goal, constraints, and acceptance metrics.

  2. Keep one verifiable goal per stage; have the model run commands and write their output to a log.

  3. Let benchmark feedback drive the next round; do not substitute “it looks faster” for test results.

  4. Run independent verification before final delivery, checking the code, charts, experiment configuration, and raw logs.

  5. Set human handoff points, a cost ceiling, and stopping conditions for long tasks to avoid feedback-free loops.

Original evidence and data

  • MiniMax's official description: M3 autonomously reproduced an ICLR paper for nearly 12 hours, producing 18 commits and 23 experiment charts and completing the core experiments.

  • In a CUDA optimization case, it completed 147 benchmark submissions and 1,959 tool calls in about 24 hours; FP8 Hopper peak utilization rose from 7.6% to 71.3%, a 9.4× speedup over the initial version.

  • MiniMax Code's Agent Team uses an adversarial Producer + Verifier loop for continuous generation, reflection, and correction.

Applicability boundaries

  • The official source does not disclose the complete system prompt, hardware configuration, all intermediate results, or failure samples; this template is a reusable transcription based on the public workflow structure, not the original prompt.

  • Results depend on clear feedback signals, tool permissions, and the harness; the same long-running performance should not be expected without a testing loop.

  • “Running autonomously for several hours” is not a quality guarantee; isolation, a cost ceiling, and human acceptance are still required.

Source excerpt or observation (short quote for compliance only)

  • The official source expands the focus of next-generation Agent coding from one-off code generation to “long-term collaboration capability.”

Source and dates

MiniMax official blog · Source date: 2026-06-01 · Edited: 2026-09-20

Read the original source
Variable checklist

Still to replace: 4

{{GOAL}}{{MATERIALS}}{{SUCCESS_CHECKS}}{{BUDGET}}

Related prompts

MiniMax M3: Reddit: Caching, Context, and Billing Verification in an M3 Agent Prompt WorkflowMiniMax M3: MiniMax Official: M-Series Prompting Best PracticesMiniMax M3: Google Supplement: Integration Prompting for Official MiniMax M3 with Claude Code / OpenCodeMiniMax M3: Reddit: MiniMax-M3 Routing and Orchestration for Long Tasks in Claude Code

Related reviews

MiniMax M3: X: FutureX real-time forecasting leaderboard — MiniMax-M3-based agent in seventh placeMiniMax M3: Google supplement: Artificial Analysis's public metrics for MiniMax-M3MiniMax M3: Official MiniMax M3 release: coding benchmarks, long context, and real long-task casesMiniMax M3: Reddit: Real-Project Benchmark — MiniMax-M3, MiMo 2.5 Pro, and Kimi K2.6

Read the full analysis

Overview · English

MiniMax M3: 1M Context, Coding Power, and the Quota Catch

A source-led MiniMax M3 overview covering M2.7 changes, API and Token Plan access, provider costs, workload fit, Tabbit boundaries, and unknowns.

MiniMax M3

Use MiniMax M3 in Tabbit

Run this guide in the environment listed above. Downloading does not transfer the template or establish model availability for your account.