Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiMo-V2.6-Flash · Media / benchmark · Vendor report

MiMo-V2.6-Flash Official Benchmarks: 30 RL Steps and Agent Results

The vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.

Media / benchmarkVendor reportEdited 2026-09-22

Test conditions

Source-specific observation
The vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.
Published conditions
These results cannot be used to infer performance across all real-world businesses, different prompts, different toolchains, or different inference parameters; the page does not provide the sample size for each benchmark, complete prompts, decoding settings, random seeds, hardware configuration。

Key data and applicable tasks

One-sentence takeaway

The vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.

Use cases

  • Tasks this can help assess: Officially published baselines for code agents, general agents, cybersecurity agents, visual agents, and long-horizon software engineering tasks.

  • Tasks this should not be extrapolated to: These results cannot be used to infer performance across all real-world businesses, different prompts, different toolchains, or different inference parameters; the page does not provide the sample size for each benchmark, complete prompts, decoding settings, random seeds, hardware configuration, or per-benchmark harness mappings.

  • Applicable model version: MiMo-V2.6-Flash; the API model name should be lowercase mimo-v2.6-flash.

  • Test environment or client: The RL training and Agent benchmarks disclosed on Xiaomi MiMo's official release page; the specific evaluation client, hardware, API endpoint, and runtime configuration are not specified.

  • Inference tier and parameters: Not specified.

Evaluation method

The official page states that MiMo-V2.6-Flash and MiMo-V2.6-Pro each completed 30 training steps in fewer than 6 days, with approximately 750,000 trajectories in total; training costs were approximately $850,000 for Flash and $2,620,000 for Pro. The page also reports relative improvements of 25% for Flash and 12% for Pro in the average pass rate on training tasks.

Visible information about training scale and configuration:

  • Each update uses 1,568 samples and supports a 1M context length; each training step contains 3.5–3.7B tokens.

  • Tasks cover Code, General, Visual, and Cyber, and the page says multi-task training was conducted through multiple harnesses.

  • The open-source end-to-end RL framework components listed on the page are verl, uni-agent, and mini-swe-agent; it also says composable lightweight mini-harnesses are provided, but does not give the harness name or configuration corresponding to each benchmark.

  • To suppress expert load drift, the MoE Router was frozen during training; reward reliability measures included reward design, adversarial evaluation, anomaly detection, and cross-checking by verifiers.

  • Training supports mixed multi-agent tasks and uses a unified trajectory representation, penalty mechanisms, and decoupling of the control plane from the data plane.

Key results

Flash official benchmark table

The table below transcribes the visible values in the original chart. The score units and percentage definitions are determined by each benchmark; the original release page does not provide a unified explanation. — indicates that the chart did not provide a value.

CategoryBenchmarkMiMo-V2.6-FlashMiMo-V2.6-ProMiMo-V2.5-ProClaude Opus 5GPT-5.6 SolFable 5
Code AgentDeepSWE v1.167.971.919.074.073.070.0
Code AgentProgramBench26.026.512.537.025.033.0
Code AgentMiMo Code Bench61.263.240.468.659.3—
General AgentAutomationBench v1.0.652.353.116.050.345.846.2
General AgentToolathlon-Verified73.676.949.180.674.977.9
General AgentGDPval-AA 2.1—16731107170815881595
General AgentAgents’ Last Exam27.631.613.231.630.825.7
General AgentTerminal Bench 4.028.834.91.549.039.942.4
General AgentTerminal Bench 2.187.689.965.289.188.884.3
General AgentOSWorld-Verified80.882.061.5†83.483.086.0
General AgentJobBench61.262.025.065.745.457.4
CybersecurityCyberGym95.194.040.0———
CybersecurityMiMo Cyber Bench77.280.20.0———
CybersecurityExploitGym6.017.80.222.130.328.4
CybersecurityExploitBench25.347.916.670.078.578.0
CybersecuritySEC Bench Pro47.566.317.7—79.1—
Visual AgentMiMo Visual Coding71.572.3—70.073.469.1

† The chart footnote states that this MiMo-V2.5 result came from MiMo-V2.5. The original does not explain the test sample size, evaluation date, runtime parameters, confidence intervals, or whether each item in the Flash table was run with the same harness.

Before and after RL training

ModelTraining stepsTrajectoriesTraining costRelative improvement in average pass rate on training tasksDeepSWE v1.1 (before → after)Improvement stated on the page
MiMo-V2.6-Flash30Approximately 750,000 combined with ProApproximately $850,00025%48.8 → 65.7Approximately 17 points
MiMo-V2.6-Pro30Approximately 750,000 combined with FlashApproximately $2,620,00012%58.4 → 72.6Approximately 14 points

The “trajectories, costs, and improvements” here are vendor reports from the release page; the original does not provide the number of Flash-only trajectories or independent evaluation records.

Raw data

  • The page says the MiMo-V2.6 series includes two native multimodal models: Pro and Flash; this article treats only the Flash column as the target model data.

  • The page text says that after RL, Flash “comprehensively outperformed MiMo-V2.5-Pro,” but does not provide the statistical basis for that judgment in the text; the item-by-item values that can be checked are those in the chart.

  • The page describes DeepSWE v1.1 as an out-of-sample long-range software engineering evaluation benchmark; the Flash release page gives a value of 48.8 → 65.7.

  • The chart explicitly compares MiMo-V2.6-Pro, MiMo-V2.5-Pro, Claude Opus 5, GPT-5.6 Sol, and Fable 5; missing values are shown as —.

  • In its open-source description of the training environment, the page also reports improvements on all 11 benchmarks starting from the SFT baseline of MiMo-V2.6-Distill-Qwen-9B, and gives values for SWE-bench Verified, MiMo Cyber Bench, Terminal Bench 2.1, and MiMo Visual Coding as examples. This is an example of training resources for Distill-Qwen-9B, not a benchmark result for MiMo-V2.6-Flash, and must not be mixed into the table above.

Conclusions and limitations

  • Vendor claim: The public Flash table shows relatively high scores on CyberGym, Terminal Bench 2.1, OSWorld-Verified, and MiMo Visual Coding; during RL training, DeepSWE v1.1 rose from 48.8 to 65.7.

  • Independent measurement: Not provided. This article does not treat the official figures as independent reproductions.

  • Version boundary: Do not rewrite data for MiMo-V2-Flash, MiMo-V2.5-Pro, MiMo-V2.6-Pro, or MiMo-V2.6-Pro-UltraSpeed as Flash data. The page's UltraSpeed description applies to Pro, with up to 20× inference speed, and is not a Flash benchmark.

  • Comparison boundary: Different models in the chart may use different evaluation implementations or available tools; the release page does not disclose a unified harness, samples, prompts, sampling parameters, or cost basis, so it cannot support strict cross-model causal conclusions.

  • Differences between text and chart labels: The body mentions Claude Fable 5.1 and GPT-6 Astra, while the benchmark chart labels are Fable 5, GPT-5.6 Sol, and Claude Opus 5; this article records each as written and does not merge or correct the vendor's names.

  • Other visible limitations: The page does not report failure cases, variance, confidence intervals, hardware, context truncation strategy, tool versions, or safety limitations for each test; the overall statement “comprehensively outperformed” also gives no aggregation formula.

Reproduction notes

  1. Open the original and confirm that the page update date is 2026-09-22, then read the body text and benchmark chart.

  2. Use the publicly documented mimo-v2.6-flash model ID; do not replace it with Pro, V2.5-Pro, Pro UltraSpeed, or the older V2-Flash.

  3. To review the table's results, obtain the official technical report, evaluation code, task-set version, harness, prompts, sampling parameters, sample size, and hardware configuration; the release page alone is insufficient to reproduce the experiment in full.

  4. For an independent reproduction, separately record the actual API/client, tool calls, cost, and failed samples, and place them in separate columns from the vendor's report on this page; do not backfill reproduced results as official values.

What this supports

  • Officially published baselines for code agents, general agents, cybersecurity agents, visual agents, and long-horizon software engineering tasks.

What this does not support

  • These results cannot be used to infer performance across all real-world businesses, different prompts, different toolchains, or different inference parameters; the page does not provide the sample size for each benchmark, complete prompts, decoding settings, random seeds, hardware configuration。

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Xiaomi MiMo official documentation · Xiaomi MiMo (official vendor) · Original publication date 2026-09-22 · Site edit date 2026-09-22

Open original source

MiMo-V2.6-Flash

Compare MiMo-V2.6-Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

MiMo-V2.6-Flash Review: High-Throughput Automation Workhorse, Conditional Agent

A source-backed MiMo-V2.6-Flash review analyzing 15B active MoE throughput, benchmark limits, long-horizon recovery cliffs, pricing, and workload fit.

Pricing · English

MiMo-V2.6-Flash Pricing: Official Rate Card, Cache Levers, and Cost per Task

A practical decision guide to MiMo-V2.6-Flash pricing: official API rates, prompt cache economics, MoE throughput, and high-volume task budgets.

Comparison · English

MiMo-V2.6-Pro vs MiMo-V2.6-Flash: Which Xiaomi MoE Model Fits Your Workload?

A head-to-head comparison of MiMo-V2.6-Pro and Flash: 1.02T vs 309B MoE architecture, 3.1x pricing delta, reasoning token overhead, agent benchmarks, and decision matrix.

Related reviews

MiMo-V2.6-Flash-RL Hugging Face Official Benchmarks and Deployment BoundariesThe official model card defines XiaomiMiMo/MiMo-V2.6-Flash-RL as the efficiency-balanced checkpoint in the MiMo-V2.6 series and reports its results on code, general Agent, cybersecurity, and visual Agent benchmarks. However, the evaluation hardware, sample sizes, complete harnesses, prompts, and decoding settings have not been disclosed.BenchLM: Same-Family Cost and Public Benchmark Comparison of MiMo-V2.6-Flash and ProThe BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner。MiMo-V2.6-Flash Official X Release Thread: Flash's Benchmark Positioning and Dual-Model StrategyThe official post positions MiMo-V2.6-Flash as an omnimodal model alongside Pro, emphasizing scaled reinforcement learning, coding/general Agent/cybersecurity/visual Agent capabilities, and open reproduction; the attached image gives Flash's individual benchmark scores but does not disclose the sample size, complete harness, hardware, prompts, or decoding parameters。MiMo-V2.6-Flash: First-hand Reddit Feedback on Compiler and GC DevelopmentOne Reddit user said they had used the model, written exactly as “MiMo V2.6 flash,” for compiler/GC development without encountering any problems and planned to keep using it. However, they provided no task details, environment, prompts, parameters, sample size, or objective metrics, so this can only serve as a personal usability signal.MiMo-V2.6-Flash Web Search Tool-Calling WorkflowFor mimo-v2.6-flash, first enable the Web Search Plugin in MiMo Console, then call the web_search tool through OpenAI Chat Completions; when real-time information is needed, use force_search: true, and use max_keyword to control the number of concurrent keywords per round and potential call costs.MiMo-V2.6-Flash Deep Thinking Configuration and Multi-turn Tool-calling Workflowmimo-v2.6-flash supports toggling deep thinking with thinking.type, which is enabled by default. When it is enabled, do not customize temperature or top_p, and pass through the historical assistant messages' reasoning_content in full during multi-turn tool calls.MiMo-V2.6-Flash Structured Output: JSON Mode Configuration and Validation WorkflowThe official documentation lists mimo-v2.6-flash as a model that supports JSON mode. When calling it, set response_format={"type": "json_object"} and explicitly require the system or user message to return JSON only, with fields, hierarchy, and types fully defined. This mode guarantees only valid JSON syntax, not the business structure, so production environments should still validate against a JSON Schema.MiMo-V2.6-Flash Image Understanding Inputs and Multi-image WorkflowThe official documentation lists mimo-v2.6-flash as a supported image-understanding model. Images can be provided through a public URL or Base64, and multiple images can be compared; the documentation does not provide a Flash-specific response, so the Pro example output, token usage, and results shown on the page cannot be extrapolated to Flash.