Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiMo-V2.6-Flash · Media / benchmark · Vendor report

MiMo-V2.6-Flash-RL Hugging Face Official Benchmarks and Deployment Boundaries

The official model card defines XiaomiMiMo/MiMo-V2.6-Flash-RL as the efficiency-balanced checkpoint in the MiMo-V2.6 series and reports its results on code, general Agent, cybersecurity, and visual Agent benchmarks. However, the evaluation hardware, sample sizes, complete harnesses, prompts, and decoding settings have not been disclosed.

Media / benchmarkVendor reportEdited 2026-09-22

Test conditions

Source-specific observation
The official model card defines XiaomiMiMo/MiMo-V2.6-Flash-RL as the efficiency-balanced checkpoint in the MiMo-V2.6 series and reports its results on code, general Agent, cybersecurity, and visual Agent benchmarks. However, the evaluation hardware, sample sizes, complete harnesses, prompts, and decoding settings have not been disclosed.
Published conditions
Accuracy, latency, throughput, cost, and production stability for unlisted tasks or under different harnesses or decoding parameters; cross-model comparisons in the table cannot replace independent testing under matched conditions.

Key data and applicable tasks

One-sentence takeaway

The official model card defines XiaomiMiMo/MiMo-V2.6-Flash-RL as the efficiency-balanced checkpoint in the MiMo-V2.6 series and reports its results on code, general Agent, cybersecurity, and visual Agent benchmarks. However, the evaluation hardware, sample sizes, complete harnesses, prompts, and decoding settings have not been disclosed. The results should therefore be treated as vendor-reported figures and cannot be directly extrapolated to all tasks or independently reproduced.

Use cases

  • Tasks suitable for assessment: Official capability positioning for long-context, multimodal-input, code Agent, general Agent, cybersecurity, and visual coding tasks; local deployment boundaries with SGLang/vLLM.

  • Tasks unsuitable for extrapolation: Accuracy, latency, throughput, cost, and production stability for unlisted tasks or under different harnesses or decoding parameters; cross-model comparisons in the table cannot replace independent testing under matched conditions.

  • Applicable model version: Only MiMo-V2.6-Flash-RL; MiMo-V2.6 Flash in the table corresponds to this checkpoint. This article does not fold data from MiMo-V2.6-Pro-RL, MiMo-V2.5-Pro, or other older/Pro variants into Flash.

  • Test environment or client: Hugging Face model card; visible deployment examples include SGLang, vLLM, Transformers, and Docker. The evaluation hardware, operating system, drivers, service versions, and API client are not specified.

  • Inference settings and parameters: The model card recommends temperature=1.0 and top_p=0.95; it does not specify whether each benchmark used these settings. The SGLang example uses --tp 8 --dp 2, --mem-fraction-static 0.65, EAGLE draft parameters, and --reasoning-parser mimo --tool-call-parser mimo; the vLLM example uses --tensor-parallel-size 4, --gpu-memory-utilization 0.95, --max-model-len auto, and the same parser configuration.

Evaluation method

The method information visible in the model card falls into two parts: a training/alignment description and a results table:

  1. Training/alignment description: The official description says that one mixed RL run covered code, general Agent, vision, and cybersecurity; tasks and multiple harnesses were mixed in the same batch. The visible scale for asynchronous GRPO is 1,568 prompts × 16 rollouts per step, with each update involving billions of tokens. GRS generates task-oriented rubrics from within-group comparative rollouts, while GAR ranks successful trajectories and reallocates advantage.

  2. Results table: The model card lists four benchmark groups - code Agent, general Agent, cybersecurity, and visual Agent - and places Flash alongside MiMo-V2.6 Pro, MiMo-V2.5 Pro, Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5.

  3. Visible evaluation conditions: Some benchmarks include version numbers, such as DeepSWE v1.1, AutomationBench v1.0.6, GDPval-AA 2.1, and Terminal Bench 4.0/2.1; versions for the others, sample sizes, prompts, few-shot settings, tool configurations, scoring scripts, hardware, and repeat counts are not specified.

  4. Nature of the results: All figures below are vendor-reported values from the model card; the page provides no independent retest record for Flash or itemized evaluation commands that can be reproduced end to end.

Key results

All columns are transcribed from the model card's original table (- retained as the original missing marker):

CategoryBenchmark (version)MiMo-V2.6 ProMiMo-V2.6 FlashMiMo-V2.5 ProClaude Opus 5GPT-5.6 SolClaude Fable 5
Code AgentDeepSWE v1.171.967.919.074.073.070.0
Code AgentProgramBench26.526.012.537.025.033.0
Code AgentMiMo Code Bench63.261.240.468.659.3-
General AgentAutomationBench v1.0.653.152.316.050.345.846.2
General AgentToolathlon-Verified76.973.649.180.674.977.9
General AgentGDPval-AA 2.11673-1107170815881595
General AgentAgents’ Last Exam31.627.613.231.630.825.7
General AgentTerminal Bench 4.034.928.81.549.039.942.4
General AgentTerminal Bench 2.189.987.665.289.188.884.3
General AgentOSWorld-Verified82.080.8-83.483.086.0
General AgentJobBench62.061.225.065.745.457.4
CybersecurityCyberGym94.095.140.0---
CybersecurityMiMo Cyber Bench80.277.20.0---
CybersecurityExploitGym17.86.00.222.130.328.4
CybersecurityExploitBench47.925.316.670.078.578.0
CybersecuritySEC Bench Pro66.347.517.7-79.1-
Visual AgentMiMo VisualCoding72.371.5-70.073.469.1

The - entries in the table are missing-value markers from the original and should not be interpreted as 0. Because the evaluation conditions for each column are not specified, the table cannot establish that the different models used the same harness, prompts, or time window.

Raw data

Target checkpoint and capability summary

  • Exact repository identifier: XiaomiMiMo/MiMo-V2.6-Flash-RL; both the model card title and deployment commands use this identifier.

  • Positioning within the series: The efficiency-balanced checkpoint in the MiMo-V2.6 series.

  • Architecture: Sparse MoE; the model card summary gives 309B total parameters and 15B activated parameters.

  • Context: Maximum context of 1M tokens; the repository's config.json sets max_position_embeddings to 1048576.

  • Modalities: Text, image, video, and audio; tags include multimodal, vision-language, audio, video-understanding, agent, and long-context.

  • Vision encoder: 681M-parameter MiMo ViT with 28 layers (24 SWA + 4 full-attention).

  • Audio encoder: 308M-parameter AudioTokenizer + 127M audio patch encoder.

  • Backbone configuration: 48 layers (39 SWA + 9 GA), hidden size 4096, SWA Q/KV heads 64/8, GA Q/KV heads 64/4, sliding window 128, 256 routed experts, and 8 experts activated per token.

  • Speculative decoding: A 5-layer SWA MTP speculative decoder; the model card says that each forward pass predicts the next 7 tokens, while the deployment examples additionally configure EAGLE draft steps.

Parameter summary discrepancy

The same Hugging Face page's sidebar separately shows Model size: 159B params and lists F32, BF16, F8_E4M3, and U8 tensor types. The model card body’s 309B total / 15B activated is an architecture summary. The page does not explain how 159B corresponds to those figures, nor whether 159B is calculated according to a particular weight format, with modules removed, or under a different counting convention. Therefore, this article does not force a conversion or consolidation:

  • Vendor architecture convention: 309B total parameters and 15B activated parameters; used to describe the sparse MoE computation structure.

  • Hugging Face page summary convention: 159B params; recorded only as page metadata and not a substitute for the architecture figures in the body.

  • Reproduction recommendation: When citing parameter counts, state the convention used and further verify against the repository's current config.json, weight index, and actual loading logs; do not interpret 159B as 15B activated parameters, and do not merge Flash with Pro.

Conclusions and limitations

  • In the official table, Flash performs relatively strongly on CyberGym (95.1) and Terminal Bench 2.1 (87.6), among others; on ExploitGym (6.0), ExploitBench (25.3), and SEC Bench Pro (47.5), it is markedly below MiMo-V2.6 Pro in the same table. These are item-by-item results and do not constitute an overall ranking.

  • The coverage of the results is limited: the Flash value for GDPval-AA 2.1 is blank, and the sample sizes, hardware, prompts, tools/harnesses, and scoring details for each benchmark have not been disclosed; latency, throughput, cost, or real-world business success rates cannot be inferred from them.

  • The model card defines a local deployment boundary, but it is not a performance measurement: the SGLang example depends on trust-remote-code, TP/DP, and MTP/EAGLE configurations; the vLLM example notes that the stable release may lag behind and points to the prebuilt image vllm/vllm-openai:mimov25-cu129. Actual compatibility should be verified against the current framework version, available VRAM, and weight format.

  • The cross-model columns (Claude, GPT, and others) are comparison values in the vendor's table; their sources, dates, and harness conditions are not specified, so they should not be treated as measurements independently conducted under matched conditions.

Reproduction notes

  1. Download or mount the exact repository XiaomiMiMo/MiMo-V2.6-Flash-RL; do not replace it with Pro, V2.5, or another Flash version.

  2. Deploy according to the model card's Transformers, SGLang, or vLLM examples; retain the trust_remote_code, multimodal encoder, and reasoning/tool-call parser configurations, and record the framework, image, driver, GPU, VRAM, parallelism, and weight format.

  3. Use temperature=1.0 and top_p=0.95 as the sampling starting point given by the model card; record the actual max tokens, tool settings, prompts, sample sizes, and scoring scripts for each item. The model card does not provide these evaluation details, so reproducers must fill them in themselves and cannot claim matched conditions with the official figures.

  4. Run benchmarks at the same versions and save the raw outputs, scoring results, and failed samples separately; keep - in the table as “not reported” and do not rewrite it as 0.

  5. For the parameter discrepancy, record the model card's 309B/15B architecture summary, the Hugging Face 159B page summary, config.json, and actual loading logs together; state the counting convention before making comparisons.

What this supports

  • Official capability positioning for long-context, multimodal-input, code Agent, general Agent, cybersecurity, and visual coding tasks; local deployment boundaries with SGLang/vLLM.

What this does not support

  • Accuracy, latency, throughput, cost, and production stability for unlisted tasks or under different harnesses or decoding parameters; cross-model comparisons in the table cannot replace independent testing under matched conditions.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Hugging Face · Xiaomi MiMo Team · Original publication date Unknown · Site edit date 2026-09-22

Open original source

MiMo-V2.6-Flash

Compare MiMo-V2.6-Flash in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

MiMo-V2.6-Flash Review: High-Throughput Automation Workhorse, Conditional Agent

A source-backed MiMo-V2.6-Flash review analyzing 15B active MoE throughput, benchmark limits, long-horizon recovery cliffs, pricing, and workload fit.

Pricing · English

MiMo-V2.6-Flash Pricing: Official Rate Card, Cache Levers, and Cost per Task

A practical decision guide to MiMo-V2.6-Flash pricing: official API rates, prompt cache economics, MoE throughput, and high-volume task budgets.

Comparison · English

MiMo-V2.6-Pro vs MiMo-V2.6-Flash: Which Xiaomi MoE Model Fits Your Workload?

A head-to-head comparison of MiMo-V2.6-Pro and Flash: 1.02T vs 309B MoE architecture, 3.1x pricing delta, reasoning token overhead, agent benchmarks, and decision matrix.

Related reviews

MiMo-V2.6-Flash: First-hand Reddit Feedback on Compiler and GC DevelopmentOne Reddit user said they had used the model, written exactly as “MiMo V2.6 flash,” for compiler/GC development without encountering any problems and planned to keep using it. However, they provided no task details, environment, prompts, parameters, sample size, or objective metrics, so this can only serve as a personal usability signal.MiMo-V2.6-Flash Official Benchmarks: 30 RL Steps and Agent ResultsThe vendor reports that after 30 RL steps and approximately 750,000 cumulative trajectories, MiMo-V2.6-Flash improved from 48.8 to 65.7 on DeepSWE v1.1; in the official Agent benchmark table, Flash scored 95.1 on CyberGym, 87.6 on Terminal Bench 2.1, and 71.5 on MiMo Visual Coding. All of these are results disclosed on Xiaomi's page, not independent reproductions.BenchLM: Same-Family Cost and Public Benchmark Comparison of MiMo-V2.6-Flash and ProThe BenchLM page lists 12 shared public benchmark results and estimated costs for three fixed-token scenarios for MiMo-V2.6-Flash and MiMo-V2.6-Pro: Flash wins one of the 12 results, CyberGym, while Pro wins the other 11; however, neither model is ranked in the current public ranking channels, and the page explicitly does not name an overall quality winner。MiMo-V2.6-Flash Official X Release Thread: Flash's Benchmark Positioning and Dual-Model StrategyThe official post positions MiMo-V2.6-Flash as an omnimodal model alongside Pro, emphasizing scaled reinforcement learning, coding/general Agent/cybersecurity/visual Agent capabilities, and open reproduction; the attached image gives Flash's individual benchmark scores but does not disclose the sample size, complete harness, hardware, prompts, or decoding parameters。MiMo-V2.6-Flash Web Search Tool-Calling WorkflowFor mimo-v2.6-flash, first enable the Web Search Plugin in MiMo Console, then call the web_search tool through OpenAI Chat Completions; when real-time information is needed, use force_search: true, and use max_keyword to control the number of concurrent keywords per round and potential call costs.MiMo-V2.6-Flash Deep Thinking Configuration and Multi-turn Tool-calling Workflowmimo-v2.6-flash supports toggling deep thinking with thinking.type, which is enabled by default. When it is enabled, do not customize temperature or top_p, and pass through the historical assistant messages' reasoning_content in full during multi-turn tool calls.MiMo-V2.6-Flash Structured Output: JSON Mode Configuration and Validation WorkflowThe official documentation lists mimo-v2.6-flash as a model that supports JSON mode. When calling it, set response_format={"type": "json_object"} and explicitly require the system or user message to return JSON only, with fields, hierarchy, and types fully defined. This mode guarantees only valid JSON syntax, not the business structure, so production environments should still validate against a JSON Schema.MiMo-V2.6-Flash Image Understanding Inputs and Multi-image WorkflowThe official documentation lists mimo-v2.6-flash as a supported image-understanding model. Images can be provided through a public URL or Base64, and multiple images can be compared; the documentation does not provide a Flash-specific response, so the Pro example output, token usage, and results shown on the page cannot be extrapolated to Flash.