MiMo-V2.6-Flash · Media / benchmark · Vendor report
The vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.
The vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.
Tasks this can help assess: Officially published baselines for code agents, general agents, cybersecurity agents, visual agents, and long-horizon software engineering tasks.
Tasks this should not be extrapolated to: These results cannot be used to infer performance across all real-world businesses, different prompts, different toolchains, or different inference parameters; the page does not provide the sample size for each benchmark, complete prompts, decoding settings, random seeds, hardware configuration, or per-benchmark harness mappings.
Applicable model version: MiMo-V2.6-Flash; the API model name should be lowercase mimo-v2.6-flash.
Test environment or client: The RL training and Agent benchmarks disclosed on Xiaomi MiMo's official release page; the specific evaluation client, hardware, API endpoint, and runtime configuration are not specified.
Inference tier and parameters: Not specified.
The official page states that MiMo-V2.6-Flash and MiMo-V2.6-Pro each completed 30 training steps in fewer than 6 days, with approximately 750,000 trajectories in total; training costs were approximately $850,000 for Flash and $2,620,000 for Pro. The page also reports relative improvements of 25% for Flash and 12% for Pro in the average pass rate on training tasks.
Visible information about training scale and configuration:
Each update uses 1,568 samples and supports a 1M context length; each training step contains 3.5–3.7B tokens.
Tasks cover Code, General, Visual, and Cyber, and the page says multi-task training was conducted through multiple harnesses.
The open-source end-to-end RL framework components listed on the page are verl, uni-agent, and mini-swe-agent; it also says composable lightweight mini-harnesses are provided, but does not give the harness name or configuration corresponding to each benchmark.
To suppress expert load drift, the MoE Router was frozen during training; reward reliability measures included reward design, adversarial evaluation, anomaly detection, and cross-checking by verifiers.
Training supports mixed multi-agent tasks and uses a unified trajectory representation, penalty mechanisms, and decoupling of the control plane from the data plane.
The table below transcribes the visible values in the original chart. The score units and percentage definitions are determined by each benchmark; the original release page does not provide a unified explanation. — indicates that the chart did not provide a value.
| Category | Benchmark | MiMo-V2.6-Flash | MiMo-V2.6-Pro | MiMo-V2.5-Pro | Claude Opus 5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|---|---|---|---|
| Code Agent | DeepSWE v1.1 | 67.9 | 71.9 | 19.0 | 74.0 | 73.0 | 70.0 |
| Code Agent | ProgramBench | 26.0 | 26.5 | 12.5 | 37.0 | 25.0 | 33.0 |
| Code Agent | MiMo Code Bench | 61.2 | 63.2 | 40.4 | 68.6 | 59.3 | — |
| General Agent | AutomationBench v1.0.6 | 52.3 | 53.1 | 16.0 | 50.3 | 45.8 | 46.2 |
| General Agent | Toolathlon-Verified | 73.6 | 76.9 | 49.1 | 80.6 | 74.9 | 77.9 |
| General Agent | GDPval-AA 2.1 | — | 1673 | 1107 | 1708 | 1588 | 1595 |
| General Agent | Agents’ Last Exam | 27.6 | 31.6 | 13.2 | 31.6 | 30.8 | 25.7 |
| General Agent | Terminal Bench 4.0 | 28.8 | 34.9 | 1.5 | 49.0 | 39.9 | 42.4 |
| General Agent | Terminal Bench 2.1 | 87.6 | 89.9 | 65.2 | 89.1 | 88.8 | 84.3 |
| General Agent | OSWorld-Verified | 80.8 | 82.0 | 61.5† | 83.4 | 83.0 | 86.0 |
| General Agent | JobBench | 61.2 | 62.0 | 25.0 | 65.7 | 45.4 | 57.4 |
| Cybersecurity | CyberGym | 95.1 | 94.0 | 40.0 | — | — | — |
| Cybersecurity | MiMo Cyber Bench | 77.2 | 80.2 | 0.0 | — | — | — |
| Cybersecurity | ExploitGym | 6.0 | 17.8 | 0.2 | 22.1 | 30.3 | 28.4 |
| Cybersecurity | ExploitBench | 25.3 | 47.9 | 16.6 | 70.0 | 78.5 | 78.0 |
| Cybersecurity | SEC Bench Pro | 47.5 | 66.3 | 17.7 | — | 79.1 | — |
| Visual Agent | MiMo Visual Coding | 71.5 | 72.3 | — | 70.0 | 73.4 | 69.1 |
† The chart footnote states that this MiMo-V2.5 result came from MiMo-V2.5. The original does not explain the test sample size, evaluation date, runtime parameters, confidence intervals, or whether each item in the Flash table was run with the same harness.
| Model | Training steps | Trajectories | Training cost | Relative improvement in average pass rate on training tasks | DeepSWE v1.1 (before → after) | Improvement stated on the page |
|---|---|---|---|---|---|---|
| MiMo-V2.6-Flash | 30 | Approximately 750,000 combined with Pro | Approximately $850,000 | 25% | 48.8 → 65.7 | Approximately 17 points |
| MiMo-V2.6-Pro | 30 | Approximately 750,000 combined with Flash | Approximately $2,620,000 | 12% | 58.4 → 72.6 | Approximately 14 points |
The “trajectories, costs, and improvements” here are vendor reports from the release page; the original does not provide the number of Flash-only trajectories or independent evaluation records.
The page says the MiMo-V2.6 series includes two native multimodal models: Pro and Flash; this article treats only the Flash column as the target model data.
The page text says that after RL, Flash “comprehensively outperformed MiMo-V2.5-Pro,” but does not provide the statistical basis for that judgment in the text; the item-by-item values that can be checked are those in the chart.
The page describes DeepSWE v1.1 as an out-of-sample long-range software engineering evaluation benchmark; the Flash release page gives a value of 48.8 → 65.7.
The chart explicitly compares MiMo-V2.6-Pro, MiMo-V2.5-Pro, Claude Opus 5, GPT-5.6 Sol, and Fable 5; missing values are shown as —.
In its open-source description of the training environment, the page also reports improvements on all 11 benchmarks starting from the SFT baseline of MiMo-V2.6-Distill-Qwen-9B, and gives values for SWE-bench Verified, MiMo Cyber Bench, Terminal Bench 2.1, and MiMo Visual Coding as examples. This is an example of training resources for Distill-Qwen-9B, not a benchmark result for MiMo-V2.6-Flash, and must not be mixed into the table above.
Vendor claim: The public Flash table shows relatively high scores on CyberGym, Terminal Bench 2.1, OSWorld-Verified, and MiMo Visual Coding; during RL training, DeepSWE v1.1 rose from 48.8 to 65.7.
Independent measurement: Not provided. This article does not treat the official figures as independent reproductions.
Version boundary: Do not rewrite data for MiMo-V2-Flash, MiMo-V2.5-Pro, MiMo-V2.6-Pro, or MiMo-V2.6-Pro-UltraSpeed as Flash data. The page's UltraSpeed description applies to Pro, with up to 20× inference speed, and is not a Flash benchmark.
Comparison boundary: Different models in the chart may use different evaluation implementations or available tools; the release page does not disclose a unified harness, samples, prompts, sampling parameters, or cost basis, so it cannot support strict cross-model causal conclusions.
Differences between text and chart labels: The body mentions Claude Fable 5.1 and GPT-6 Astra, while the benchmark chart labels are Fable 5, GPT-5.6 Sol, and Claude Opus 5; this article records each as written and does not merge or correct the vendor's names.
Other visible limitations: The page does not report failure cases, variance, confidence intervals, hardware, context truncation strategy, tool versions, or safety limitations for each test; the overall statement “comprehensively outperformed” also gives no aggregation formula.
Open the original and confirm that the page update date is 2026-09-22, then read the body text and benchmark chart.
Use the publicly documented mimo-v2.6-flash model ID; do not replace it with Pro, V2.5-Pro, Pro UltraSpeed, or the older V2-Flash.
To review the table's results, obtain the official technical report, evaluation code, task-set version, harness, prompts, sampling parameters, sample size, and hardware configuration; the release page alone is insufficient to reproduce the experiment in full.
For an independent reproduction, separately record the actual API/client, tool calls, cost, and failed samples, and place them in separate columns from the vendor's report on this page; do not backfill reproduced results as official values.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
Xiaomi MiMo official documentation · Xiaomi MiMo (official vendor) · Original publication date 2026-09-22 · Site edit date 2026-09-22
Open original sourceMiMo-V2.6-Flash
Download the Tabbit client to check model access