MiMo-V2.6-Flash · Media / benchmark · Editorial analysis
The BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner。
The BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner, so these results cannot be rewritten as an overall victory for Pro or Flash.
Tasks this is suitable for: Reviewing the public benchmark coverage, item-level scores, context specifications, and cost differences under fixed token usage for two models in the same family in the same BenchLM page snapshot.
Tasks this is not suitable for extrapolating to: Overall quality rankings, unlisted tasks, real-world business success rates, latency, throughput, stability, or performance under different editors, harnesses, reasoning efforts, or prompts. The page also does not provide the sample size, raw outputs, or confidence intervals for BenchLM's independent runs.
Applicable model versions: The page compares only MiMo-V2.6-Flash and MiMo-V2.6-Pro; it does not mix MiMo-V2.6-Flash-RL, older MiMo versions, MiMo-V2.5-Pro, or other models into this page's data.
Test environment or client: BenchLM.ai comparison page; the shared source for the benchmark rows is linked to the Xiaomi MiMo-V2.6 technical report. The specific API client, hardware, operating system, runner, and harness are not specified.
Reasoning tier and parameters: The page specifications for both models are Reasoning; benchmark reasoning parameters, prompts, sampling settings, and effort are not specified.
The BenchLM page snapshot was updated on 2026-09-21, and the data generation date was 2026-09-21; the method version shown by the public data interface is bench-align-v5.5-2026-09-04. The page's comparison rules are as follows:
The top-level model-card scores, public rankings, 90% intervals, and evidence statuses all show as unavailable (—, Evidence status unavailable, 90% interval unavailable); accordingly, the page states that “at least one model is not scored in the current public ranking channels” and does not name an overall quality winner.
There are 12 shared results in total, with 0 Flash-only and 0 Pro-only; “like-for-like categories” is 0/8. The 12 are counted by rows in the public results ledger; Terminal-Bench 2.1 appears once in each of the Agentic and Coding categories.
Agentic, Coding, and Knowledge use the BenchAlign lane; the remaining categories use provisional/weighted public rows. A result counts as like-for-like only when both models' scores are based on Supported evidence or the same weighted set; directional or incomparable rows do not produce a winner.
The page explains that unranked scores in the rankings can only be retained on the lane's scale and cannot be used to name a winner; under the practical rule, like-for-like rows with a difference of no more than 0.5 points are treated as ties. All eight categories on this page are Not ranked / Not comparable.
All 12 raw rows are marked Shared source and link to the same https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf. This shows that the page includes a public source; it does not mean that BenchLM disclosed an independent retest process on this page.
The page uses three fixed token mixes to convert the “listed standard API rates” into estimated per-request costs; every scenario shows “Fits in one request”.
| Scenario | Token assumption | MiMo-V2.6-Flash | MiMo-V2.6-Pro | Page reading |
|---|---|---|---|---|
| Chat turn | 1K fresh input + 500 output | $0.00028 | $0.00087 | Flash lower, listed-rates |
| Repository review | 50K fresh input + 3K output | $0.00784 | $0.02436 | Flash lower |
| Cache-heavy agent loop | 200K cached + 20K fresh input + 10K output | $0.00616 | $0.01812 | Flash lower |
The cached-input prices listed in the page's specifications section are: Flash $0.0028 / 1M cached input tokens and Pro $0.0036 / 1M cached input tokens. The comparison body does not separately show the fresh-input/output unit prices; the standard prices consistent with the three page amounts can be calculated as follows:
Flash: 1K×$0.14/M + 500×$0.28/M = $0.00028
Pro: 1K×$0.435/M + 500×$0.87/M = $0.00087
Flash cache loop = 200K×$0.0028/M + 20K×$0.14/M + 10K×$0.28/M = $0.00616
Pro cache loop = 200K×$0.0036/M + 20K×$0.435/M + 10K×$0.87/M = $0.01812The fresh-input/output unit prices above are arithmetic back-calculations from the amounts already shown on the page, not prices separately listed by the page or independently verified prices. The page specifies that when the cached-input price is missing, the estimate falls back to the published standard input price; both models on this page list a cached-input price.
The values in the page's raw results ledger are as follows. Percentages retain the page's display format; “leads” is the page's row-level label and does not represent an overall ranking.
| Category | Benchmark | MiMo-V2.6-Flash | MiMo-V2.6-Pro | Page row-level label |
|---|---|---|---|---|
| Agentic | Toolathlon-Verified | 73.6% | 76.9% | Pro leads |
| Agentic | AutomationBench | 52.3% | 53.1% | Pro leads |
| Agentic | Agents’ Last Exam | 27.6% | 31.6% | Pro leads |
| Agentic | Terminal-Bench 4.0 | 28.80% | 34.90% | Pro leads |
| Agentic | Terminal-Bench 2.1 | 87.6% | 89.9% | Pro leads |
| Agentic | OSWorld-Verified | 80.8% | 82% | Pro leads |
| Agentic | JobBench | 61.2% | 62.0% | Pro leads |
| Agentic | CyberGym | 95.1% | 94.0% | Flash leads |
| Agentic | ExploitGym | 6.0% | 17.8% | Pro leads |
| Coding | DeepSWE | 67.9% | 71.9% | Pro leads |
| Coding | ProgramBench | 26.0% | 26.5% | Pro leads |
| Coding | Terminal-Bench 2.1 | 87.6% | 89.9% | Pro leads |
| Category | BenchLM lane | Flash | Pro | Page status |
|---|---|---|---|---|
| Agentic | BenchAlign | Not ranked | Not ranked | Not comparable; 9 vs 9 public rows |
| Coding | BenchAlign | Not ranked | Not ranked | Not comparable; 3 vs 3 public rows |
| Reasoning | Provisional | Not ranked | Not ranked | Not comparable; 0 vs 0 weighted rows |
| Multimodal | Provisional | Not ranked | Not ranked | Not comparable; 0 vs 0 weighted rows |
| Knowledge | BenchAlign | Not ranked | Not ranked | Not comparable; 0 vs 0 public rows |
| Multilingual | Provisional | Not ranked | Not ranked | Not comparable; 0 vs 0 weighted rows |
| Instruction following | Provisional | Not ranked | Not ranked | Not comparable; 0 vs 0 weighted rows |
| Math | Provisional | Not ranked | Not ranked | Not comparable; 0 vs 0 weighted rows |
Specification differences between the two models (the page lists Xiaomi MiMo-V2.6 launch as the source for both):
| Specification | MiMo-V2.6-Flash | MiMo-V2.6-Pro |
|---|---|---|
| Maximum document context | 1M | 1M |
| API model ID | Not sourced | Not sourced |
| Cached-input price | $0.0028 / 1M | $0.0036 / 1M |
| Documented input | Not sourced | Not sourced |
| Documented output | Not sourced | Not sourced |
| Provider availability | Not sourced | Not sourced |
| Reasoning profile | Reasoning | Reasoning |
| Weight access | Open Weight | Open Weight |
| License | Open Weight | Open Weight |
| Release date | 2026-09-22 | 2026-09-22 |
The page notes that 1M is the maximum documented context, and the actual output-token limit may be lower; the page values for both Weight access and License are Open Weight, which should not be used to fill in a specific license name.
Page decision text: At least one model is not scored in the current public ranking channels, so the page “does not name an overall quality winner”; category rows based on Estimated evidence or different benchmark sets are only directional references and likewise do not name a winner.
Evidence counts: 12 Shared, MiMo-V2.6-Flash only 0, MiMo-V2.6-Pro only 0, Like-for-like categories 0 / 8.
Shared evidence shape: The page plots only two common results: OSWorld-Verified (Flash 80.8%, Pro 82%, normalized gap 1.2) and AutomationBench (Flash 52.3%, Pro 53.1%, normalized gap 0.8); because there are too few matched category axes, it does not generate a radar chart.
Source attributes: The API data returns attribution: Data from BenchLM.ai; the source label for every raw benchmark row is Xiaomi MiMo-V2.6 technical report.
Single-page model-card status: The score, publicRank, interval90, and evidenceStatus for both models are null; the page-visible text is respectively —, Evidence status unavailable, and 90% interval unavailable.
Page update boundary: The page was updated on 2026-09-21, and the page footer likewise says Last updated September 21, 2026; this article was collected on 2026-09-22 and is a dynamic leaderboard snapshot for that date.
Cost: Under the page's three token assumptions of 1K/500, 50K/3K, and 200K/20K/10K, Flash's estimated amount is lower than Pro's in every case; this is modeled cost based on listed-rates, not an actual bill or a throughput or latency measurement.
Item-level results: Of the 12 shared rows, only CyberGym (95.1% versus 94.0%) is marked as leading for Flash by the page; Pro leads the other 11. The repeated Terminal-Bench 2.1 row is a result of the category-ledger design and should not be counted twice as two independent experiments.
Overall-conclusion limitation: BenchLM explicitly does not name an overall quality winner for this pair; do not rewrite the row-level labels, cost advantage, or simple count across 12 rows as “Pro is stronger overall” or “Flash has better overall value for money.”
Nature of the evidence: The public table's source is the Xiaomi MiMo-V2.6 technical report, and the BenchLM page does not disclose the sample size, prompts, hardware, decoding parameters, number of repetitions, failed samples, scoring scripts, or confidence intervals for independent runs. Therefore, this article describes it as a public-source result/BenchLM page compilation, not as an independently retested result.
Comparability limitation: The page says that Agentic, Coding, and Knowledge use the BenchAlign lane, while the others use provisional/weighted rows; this page has 0/8 like-for-like, and all eight categories are incomparable. Missing category scores do not mean 0 points.
Cost limitation: The page does not separately list fresh/output prices in the body; the $0.14/$0.28 and $0.435/$0.87 back-calculated in this article are only used to explain the page amounts and cannot replace the official pricing page. The actual billing rules for cached tokens and input/output tokens, as well as context truncation, are not fully disclosed on this page.
Specification limitation: API model ID, documented input/output, and provider availability are not specified; 1M only indicates the maximum document context, and the output limit may be lower.
Version boundary: This article records only MiMo-V2.6-Flash and MiMo-V2.6-Pro. It does not bring the scores, prices, or experience of older MiMo versions, MiMo-V2.6-Flash-RL, MiMo-V2.5-Pro, or any other page or model into this page.
Open https://benchlm.ai/compare/mimo-v2-6-flash-vs-mimo-v2-6-pro and confirm the page update date, decision text, 12 shared results, and “0/8 like-for-like categories”.
Expand Browse raw public benchmark evidence, transcribe the 12 rows by Agentic/Coding category, and retain the page's percentages and the two category appearances of Terminal-Bench 2.1.
Recalculate costs using the page's three fixed token mixes; use the page-listed $0.0028/M (Flash) and $0.0036/M (Pro) for cached input. If the fresh/output unit prices back-calculated in this article are used, they must be labeled as back-calculated from the page amounts and must not be written as prices directly disclosed by the page.
For an independent retest, separately record the exact model ID, API or local deployment method, client, hardware, harness, prompts, reasoning effort, sampling parameters, sample size, scoring script, failed samples, and run date; the BenchLM page itself does not provide these fields.
Do not treat the public-source table on this page as an overall ranking of Flash and Pro, and do not treat the page's listed-rates estimate as actual production cost or an independent performance measurement.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
BenchLM.ai · Not specified (the page is attributed to BenchLM.ai data) · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceMiMo-V2.6-Flash
Download the Tabbit client to check model access