Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiMo-V2.6-Flash · Media / benchmark · Editorial analysis

BenchLM: Same-Family Cost and Public Benchmark Comparison of MiMo-V2.6-Flash and Pro

The BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner。

Media / benchmarkEditorial analysisEdited 2026-09-22

Test conditions

Source-specific observation
The BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner。
Published conditions
Overall quality rankings, unlisted tasks, real-world business success rates, latency, throughput, stability, or performance under different editors, harnesses, reasoning efforts, or prompts. The page also does not provide the sample size, raw outputs, or confidence intervals for BenchLM's independent runs.

Key data and applicable tasks

One-sentence takeaway

The BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner, so these results cannot be rewritten as an overall victory for Pro or Flash.

Use cases

  • Tasks this is suitable for: Reviewing the public benchmark coverage, item-level scores, context specifications, and cost differences under fixed token usage for two models in the same family in the same BenchLM page snapshot.

  • Tasks this is not suitable for extrapolating to: Overall quality rankings, unlisted tasks, real-world business success rates, latency, throughput, stability, or performance under different editors, harnesses, reasoning efforts, or prompts. The page also does not provide the sample size, raw outputs, or confidence intervals for BenchLM's independent runs.

  • Applicable model versions: The page compares only MiMo-V2.6-Flash and MiMo-V2.6-Pro; it does not mix MiMo-V2.6-Flash-RL, older MiMo versions, MiMo-V2.5-Pro, or other models into this page's data.

  • Test environment or client: BenchLM.ai comparison page; the shared source for the benchmark rows is linked to the Xiaomi MiMo-V2.6 technical report. The specific API client, hardware, operating system, runner, and harness are not specified.

  • Reasoning tier and parameters: The page specifications for both models are Reasoning; benchmark reasoning parameters, prompts, sampling settings, and effort are not specified.

Evaluation method

The BenchLM page snapshot was updated on 2026-09-21, and the data generation date was 2026-09-21; the method version shown by the public data interface is bench-align-v5.5-2026-09-04. The page's comparison rules are as follows:

  • The top-level model-card scores, public rankings, 90% intervals, and evidence statuses all show as unavailable (—, Evidence status unavailable, 90% interval unavailable); accordingly, the page states that “at least one model is not scored in the current public ranking channels” and does not name an overall quality winner.

  • There are 12 shared results in total, with 0 Flash-only and 0 Pro-only; “like-for-like categories” is 0/8. The 12 are counted by rows in the public results ledger; Terminal-Bench 2.1 appears once in each of the Agentic and Coding categories.

  • Agentic, Coding, and Knowledge use the BenchAlign lane; the remaining categories use provisional/weighted public rows. A result counts as like-for-like only when both models' scores are based on Supported evidence or the same weighted set; directional or incomparable rows do not produce a winner.

  • The page explains that unranked scores in the rankings can only be retained on the lane's scale and cannot be used to name a winner; under the practical rule, like-for-like rows with a difference of no more than 0.5 points are treated as ties. All eight categories on this page are Not ranked / Not comparable.

  • All 12 raw rows are marked Shared source and link to the same https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf. This shows that the page includes a public source; it does not mean that BenchLM disclosed an independent retest process on this page.

Key results

Cost scenarios

The page uses three fixed token mixes to convert the “listed standard API rates” into estimated per-request costs; every scenario shows “Fits in one request”.

ScenarioToken assumptionMiMo-V2.6-FlashMiMo-V2.6-ProPage reading
Chat turn1K fresh input + 500 output$0.00028$0.00087Flash lower, listed-rates
Repository review50K fresh input + 3K output$0.00784$0.02436Flash lower
Cache-heavy agent loop200K cached + 20K fresh input + 10K output$0.00616$0.01812Flash lower

The cached-input prices listed in the page's specifications section are: Flash $0.0028 / 1M cached input tokens and Pro $0.0036 / 1M cached input tokens. The comparison body does not separately show the fresh-input/output unit prices; the standard prices consistent with the three page amounts can be calculated as follows:

Flash: 1K×$0.14/M + 500×$0.28/M = $0.00028
Pro:   1K×$0.435/M + 500×$0.87/M = $0.00087

Flash cache loop = 200K×$0.0028/M + 20K×$0.14/M + 10K×$0.28/M = $0.00616
Pro cache loop   = 200K×$0.0036/M + 20K×$0.435/M + 10K×$0.87/M = $0.01812

The fresh-input/output unit prices above are arithmetic back-calculations from the amounts already shown on the page, not prices separately listed by the page or independently verified prices. The page specifies that when the cached-input price is missing, the estimate falls back to the published standard input price; both models on this page list a cached-input price.

Shared public benchmark rows

The values in the page's raw results ledger are as follows. Percentages retain the page's display format; “leads” is the page's row-level label and does not represent an overall ranking.

CategoryBenchmarkMiMo-V2.6-FlashMiMo-V2.6-ProPage row-level label
AgenticToolathlon-Verified73.6%76.9%Pro leads
AgenticAutomationBench52.3%53.1%Pro leads
AgenticAgents’ Last Exam27.6%31.6%Pro leads
AgenticTerminal-Bench 4.028.80%34.90%Pro leads
AgenticTerminal-Bench 2.187.6%89.9%Pro leads
AgenticOSWorld-Verified80.8%82%Pro leads
AgenticJobBench61.2%62.0%Pro leads
AgenticCyberGym95.1%94.0%Flash leads
AgenticExploitGym6.0%17.8%Pro leads
CodingDeepSWE67.9%71.9%Pro leads
CodingProgramBench26.0%26.5%Pro leads
CodingTerminal-Bench 2.187.6%89.9%Pro leads

Category aggregation status and specifications

CategoryBenchLM laneFlashProPage status
AgenticBenchAlignNot rankedNot rankedNot comparable; 9 vs 9 public rows
CodingBenchAlignNot rankedNot rankedNot comparable; 3 vs 3 public rows
ReasoningProvisionalNot rankedNot rankedNot comparable; 0 vs 0 weighted rows
MultimodalProvisionalNot rankedNot rankedNot comparable; 0 vs 0 weighted rows
KnowledgeBenchAlignNot rankedNot rankedNot comparable; 0 vs 0 public rows
MultilingualProvisionalNot rankedNot rankedNot comparable; 0 vs 0 weighted rows
Instruction followingProvisionalNot rankedNot rankedNot comparable; 0 vs 0 weighted rows
MathProvisionalNot rankedNot rankedNot comparable; 0 vs 0 weighted rows

Specification differences between the two models (the page lists Xiaomi MiMo-V2.6 launch as the source for both):

SpecificationMiMo-V2.6-FlashMiMo-V2.6-Pro
Maximum document context1M1M
API model IDNot sourcedNot sourced
Cached-input price$0.0028 / 1M$0.0036 / 1M
Documented inputNot sourcedNot sourced
Documented outputNot sourcedNot sourced
Provider availabilityNot sourcedNot sourced
Reasoning profileReasoningReasoning
Weight accessOpen WeightOpen Weight
LicenseOpen WeightOpen Weight
Release date2026-09-222026-09-22

The page notes that 1M is the maximum documented context, and the actual output-token limit may be lower; the page values for both Weight access and License are Open Weight, which should not be used to fill in a specific license name.

Raw data

  • Page decision text: At least one model is not scored in the current public ranking channels, so the page “does not name an overall quality winner”; category rows based on Estimated evidence or different benchmark sets are only directional references and likewise do not name a winner.

  • Evidence counts: 12 Shared, MiMo-V2.6-Flash only 0, MiMo-V2.6-Pro only 0, Like-for-like categories 0 / 8.

  • Shared evidence shape: The page plots only two common results: OSWorld-Verified (Flash 80.8%, Pro 82%, normalized gap 1.2) and AutomationBench (Flash 52.3%, Pro 53.1%, normalized gap 0.8); because there are too few matched category axes, it does not generate a radar chart.

  • Source attributes: The API data returns attribution: Data from BenchLM.ai; the source label for every raw benchmark row is Xiaomi MiMo-V2.6 technical report.

  • Single-page model-card status: The score, publicRank, interval90, and evidenceStatus for both models are null; the page-visible text is respectively —, Evidence status unavailable, and 90% interval unavailable.

  • Page update boundary: The page was updated on 2026-09-21, and the page footer likewise says Last updated September 21, 2026; this article was collected on 2026-09-22 and is a dynamic leaderboard snapshot for that date.

Conclusions and limitations

  • Cost: Under the page's three token assumptions of 1K/500, 50K/3K, and 200K/20K/10K, Flash's estimated amount is lower than Pro's in every case; this is modeled cost based on listed-rates, not an actual bill or a throughput or latency measurement.

  • Item-level results: Of the 12 shared rows, only CyberGym (95.1% versus 94.0%) is marked as leading for Flash by the page; Pro leads the other 11. The repeated Terminal-Bench 2.1 row is a result of the category-ledger design and should not be counted twice as two independent experiments.

  • Overall-conclusion limitation: BenchLM explicitly does not name an overall quality winner for this pair; do not rewrite the row-level labels, cost advantage, or simple count across 12 rows as “Pro is stronger overall” or “Flash has better overall value for money.”

  • Nature of the evidence: The public table's source is the Xiaomi MiMo-V2.6 technical report, and the BenchLM page does not disclose the sample size, prompts, hardware, decoding parameters, number of repetitions, failed samples, scoring scripts, or confidence intervals for independent runs. Therefore, this article describes it as a public-source result/BenchLM page compilation, not as an independently retested result.

  • Comparability limitation: The page says that Agentic, Coding, and Knowledge use the BenchAlign lane, while the others use provisional/weighted rows; this page has 0/8 like-for-like, and all eight categories are incomparable. Missing category scores do not mean 0 points.

  • Cost limitation: The page does not separately list fresh/output prices in the body; the $0.14/$0.28 and $0.435/$0.87 back-calculated in this article are only used to explain the page amounts and cannot replace the official pricing page. The actual billing rules for cached tokens and input/output tokens, as well as context truncation, are not fully disclosed on this page.

  • Specification limitation: API model ID, documented input/output, and provider availability are not specified; 1M only indicates the maximum document context, and the output limit may be lower.

  • Version boundary: This article records only MiMo-V2.6-Flash and MiMo-V2.6-Pro. It does not bring the scores, prices, or experience of older MiMo versions, MiMo-V2.6-Flash-RL, MiMo-V2.5-Pro, or any other page or model into this page.

Reproduction notes

  1. Open https://benchlm.ai/compare/mimo-v2-6-flash-vs-mimo-v2-6-pro and confirm the page update date, decision text, 12 shared results, and “0/8 like-for-like categories”.

  2. Expand Browse raw public benchmark evidence, transcribe the 12 rows by Agentic/Coding category, and retain the page's percentages and the two category appearances of Terminal-Bench 2.1.

  3. Recalculate costs using the page's three fixed token mixes; use the page-listed $0.0028/M (Flash) and $0.0036/M (Pro) for cached input. If the fresh/output unit prices back-calculated in this article are used, they must be labeled as back-calculated from the page amounts and must not be written as prices directly disclosed by the page.

  4. For an independent retest, separately record the exact model ID, API or local deployment method, client, hardware, harness, prompts, reasoning effort, sampling parameters, sample size, scoring script, failed samples, and run date; the BenchLM page itself does not provide these fields.

  5. Do not treat the public-source table on this page as an overall ranking of Flash and Pro, and do not treat the page's listed-rates estimate as actual production cost or an independent performance measurement.

What this supports

  • Reviewing the public benchmark coverage, item-level scores, context specifications, and cost differences under fixed token usage for two models in the same family in the same BenchLM page snapshot.

What this does not support

  • Overall quality rankings, unlisted tasks, real-world business success rates, latency, throughput, stability, or performance under different editors, harnesses, reasoning efforts, or prompts. The page also does not provide the sample size, raw outputs, or confidence intervals for BenchLM's independent runs.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

BenchLM.ai · Not specified (the page is attributed to BenchLM.ai data) · Original publication date 2026-09-21 · Site edit date 2026-09-22

Open original source

MiMo-V2.6-Flash

Compare MiMo-V2.6-Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

MiMo-V2.6-Flash Review: High-Throughput Automation Workhorse, Conditional Agent

A source-backed MiMo-V2.6-Flash review analyzing 15B active MoE throughput, benchmark limits, long-horizon recovery cliffs, pricing, and workload fit.

Pricing · English

MiMo-V2.6-Flash Pricing: Official Rate Card, Cache Levers, and Cost per Task

A practical decision guide to MiMo-V2.6-Flash pricing: official API rates, prompt cache economics, MoE throughput, and high-volume task budgets.

Comparison · English

MiMo-V2.6-Pro vs MiMo-V2.6-Flash: Which Xiaomi MoE Model Fits Your Workload?

A head-to-head comparison of MiMo-V2.6-Pro and Flash: 1.02T vs 309B MoE architecture, 3.1x pricing delta, reasoning token overhead, agent benchmarks, and decision matrix.

Related reviews

MiMo-V2.6-Flash Official Benchmarks: 30 RL Steps and Agent ResultsThe vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.MiMo-V2.6-Flash-RL Hugging Face Official Benchmarks and Deployment BoundariesThe official model card defines XiaomiMiMo/MiMo-V2.6-Flash-RL as the efficiency-balanced checkpoint in the MiMo-V2.6 series and reports its results on code, general Agent, cybersecurity, and visual Agent benchmarks. However, the evaluation hardware, sample sizes, complete harnesses, prompts, and decoding settings have not been disclosed.MiMo-V2.6-Flash Official X Release Thread: Flash's Benchmark Positioning and Dual-Model StrategyThe official post positions MiMo-V2.6-Flash as an omnimodal model alongside Pro, emphasizing scaled reinforcement learning, coding/general Agent/cybersecurity/visual Agent capabilities, and open reproduction; the attached image gives Flash's individual benchmark scores but does not disclose the sample size, complete harness, hardware, prompts, or decoding parameters。MiMo-V2.6-Flash: First-hand Reddit Feedback on Compiler and GC DevelopmentOne Reddit user said they had used the model, written exactly as “MiMo V2.6 flash,” for compiler/GC development without encountering any problems and planned to keep using it. However, they provided no task details, environment, prompts, parameters, sample size, or objective metrics, so this can only serve as a personal usability signal.MiMo-V2.6-Flash Web Search Tool-Calling WorkflowFor mimo-v2.6-flash, first enable the Web Search Plugin in MiMo Console, then call the web_search tool through OpenAI Chat Completions; when real-time information is needed, use force_search: true, and use max_keyword to control the number of concurrent keywords per round and potential call costs.MiMo-V2.6-Flash Deep Thinking Configuration and Multi-turn Tool-calling Workflowmimo-v2.6-flash supports toggling deep thinking with thinking.type, which is enabled by default. When it is enabled, do not customize temperature or top_p, and pass through the historical assistant messages' reasoning_content in full during multi-turn tool calls.MiMo-V2.6-Flash Structured Output: JSON Mode Configuration and Validation WorkflowThe official documentation lists mimo-v2.6-flash as a model that supports JSON mode. When calling it, set response_format={"type": "json_object"} and explicitly require the system or user message to return JSON only, with fields, hierarchy, and types fully defined. This mode guarantees only valid JSON syntax, not the business structure, so production environments should still validate against a JSON Schema.MiMo-V2.6-Flash Image Understanding Inputs and Multi-image WorkflowThe official documentation lists mimo-v2.6-flash as a supported image-understanding model. Images can be provided through a public URL or Base64, and multiple images can be compared; the documentation does not provide a Flash-specific response, so the Pro example output, token usage, and results shown on the page cannot be extrapolated to Flash.