MiMo-V2.6-Flash · Community source · Vendor report
The official post positions MiMo-V2.6-Flash as an omnimodal model alongside Pro, emphasizing scaled reinforcement learning, coding/general Agent/cybersecurity/visual Agent capabilities, and open reproduction; the attached image gives Flash's individual benchmark scores but does not disclose the sample size, complete harness, hardware, prompts, or decoding parameters。
The official post positions MiMo-V2.6-Flash as an omnimodal model alongside Pro, emphasizing scaled reinforcement learning, coding/general Agent/cybersecurity/visual Agent capabilities, and open reproduction; the attached image gives Flash's individual benchmark scores but does not disclose the sample size, complete harness, hardware, prompts, or decoding parameters, so these figures can only be treated as a vendor baseline.
Tasks suitable for assessment: Official capability positioning for coding agents, general agents, cybersecurity agents, and visual agents; assessing the product division between Flash and Pro and their open-source direction.
Tasks unsuitable for extrapolation: Real-world business success rates, latency, throughput, cost, stability, or performance under different toolchains or inference parameters; Pro-exclusive claims must not be attributed to Flash.
Applicable model version: MiMo-V2.6-Flash. This post does not use the checkpoint name MiMo-V2.6-Flash-RL.
Test environment or client: Xiaomi MiMo's official X post and its attached image; the evaluation hardware, client, API endpoint, and harness are not specified.
Inference tier and parameters: Not specified.
The only visible evaluation material is a table in the image attached to the original post (captioned “Table 3 Comparison of MiMo-V2.6 with previous-generation and frontier models on agentic benchmarks”). The table groups benchmarks into Code Agent, General Agent, Cybersecurity, and Visual Agent, and lists MiMo-V2.6 Pro, MiMo-V2.6 Flash, MiMo-V2.5 Pro, Claude Opus 5, GPT-5.6 Sol, and Fable 5 across the columns.
The original post says that the two models were advanced through scaled reinforcement learning and claims stronger coding, computer use, 3D reasoning, and creative capabilities; the visible thread also says that API pricing for both Pro and Flash remains unchanged from V2.5, and that Pro, Flash, the technical report, 7K+ RL task environments, an end-to-end RL framework, and composable mini-harnesses are being open-sourced. These are vendor release statements, not a complete disclosure of the evaluation method.
The source does not specify the sample size for each benchmark, task-sampling rules, evaluation date, complete prompts, few-shot setup, tool permissions, decoding parameters, hardware, number of repetitions, scoring scripts, confidence intervals, or failed samples.
The table below transcribes only the MiMo-V2.6 Flash column from the attached image; — means that no value was provided in the image.
| Category | Benchmark | MiMo-V2.6-Flash |
|---|---|---|
| Code Agent | DeepSWE v1.1 | 67.9 |
| Code Agent | ProgramBench | 26.0 |
| Code Agent | MiMo Code Bench | 61.2 |
| General Agent | AutomationBench v1.0.6 | 52.3 |
| General Agent | Toolathlon-Verified | 73.6 |
| General Agent | GDPval-AA 2.1 | — |
| General Agent | Agents’ Last Exam | 27.6 |
| General Agent | Terminal Bench 4.0 | 28.8 |
| General Agent | Terminal Bench 2.1 | 87.6 |
| General Agent | OSWorld-Verified | 80.8 |
| General Agent | JobBench | 61.2 |
| Cybersecurity | CyberGym | 95.1 |
| Cybersecurity | MiMo Cyber Bench | 77.2 |
| Cybersecurity | ExploitGym | 6.0 |
| Cybersecurity | ExploitBench | 25.3 |
| Cybersecurity | SEC Bench Pro | 47.5 |
| Visual Agent | MiMo Visual Coding | 71.5 |
Original post positioning: Introducing Xiaomi MiMo-V2.6 — Pro & Flash.; Two omnimodal models, advancing through scaled reinforcement learning.
Original post capability statement: Stronger coding, computer use, 3D reasoning and creative capabilities.
Original post open-source claim: Open model weights, technical report, RL environments and training code.
Flash-related statements in the visible thread: API pricing unchanged from V2.5, for both Pro and Flash; We’re open-sourcing Pro and Flash ... 7K+ RL task environments ... composable mini-harnesses.
Image attached to the original post: https://pbs.twimg.com/media/HSxK_3caUAAsfBl?format=jpg&name=medium. In the image's Flash column, CyberGym is 95.1, Terminal Bench 2.1 is 87.6, and MiMo Visual Coding is 71.5; the remaining items in the table above are also listed in full.
Explicit version boundary: In the original post, Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and Pro scores 46 ... both clearly refer to Pro, not Flash results.
Vendor claim: Flash was released as an omnimodal model alongside Pro, with a single Agent benchmark table showing its coding, general Agent, cybersecurity, and visual Agent scores; the official thread also emphasizes unchanged API pricing and open RL resources.
Independent measurement: None provided. This page contains no independent rerun, third-party harness, or personal experience data.
Data interpretation limits: The table is a snapshot from one official release and is insufficient to infer overall rankings or average cross-task capability. The score definitions for different benchmarks are also not uniformly explained.
Comparison limits: The comparison models and Flash in the attached image may not have used the same time period, hardware, tools, prompts, or harness; differences in the table cannot be interpreted as strict causal advantages.
Model boundary: Do not rewrite data from MiMo-V2-Flash, MiMo-V2.5-Pro, MiMo-V2.6-Pro, or MiMo-V2.6-Pro-UltraSpeed as Flash data; this page records only the MiMo-V2.6 Flash column shown in the image.
Open the original X post, confirm that the author is @XiaomiMiMo, and inspect the MiMo-V2.6 Flash column in the image attached to the original post.
Transcribe the scores item by item by category and benchmark name; keep the image's — as “not provided” and do not rewrite it as 0.
To conduct an independent rerun, obtain the task-set version, complete harness, prompts, tool configuration, decoding parameters, hardware, sample size, and scoring scripts separately; the original post itself is insufficient to reproduce the experiment.
Store independent results in a separate column from the vendor values in this article, and record the actual model identifier, client, run time, and failed samples; do not use independent results to replace the official table.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X (Twitter) · Xiaomi MiMo (@XiaomiMiMo, the vendor's official account) · Original publication date 2026-09-21 · Site edit date 2026-09-22
Open original sourceMiMo-V2.6-Flash
Download the Tabbit client to check model access