MiMo-V2.6-Pro · Community source · Vendor report
This official thread positions MiMo-V2.6-Pro as an openly built native omnimodal agent model focused on coding, computer use, 3D, design, research, and tool workflows; its rankings and examples help identify promising task directions, but they remain vendor-reported and cannot replace an independent rerun under the same harness.
This official thread positions MiMo-V2.6-Pro as an openly built native omnimodal agent model focused on coding, computer use, 3D, design, research, and tool workflows; its rankings and examples help identify promising task directions, but they remain vendor-reported and cannot replace an independent rerun under the same harness.
Suitable tasks: Long-horizon software engineering and tool use; 3D scenes driven by images, video, or text; Blender modeling; desktop search, editing, and data processing; visually guided embodied simulation; front-end, presentation, SVG, video, and music creation; materials research; and Lean 4 formalization.
Unsuitable tasks: Production decisions based only on claims of being “comparable” or on price ratios in this thread; rigorous comparisons requiring per-item inputs, complete tool schemas, random seeds, failure logs, and statistical significance; or treating a single official case as a stable success rate.
Applicable model versions: mimo-v2.6-pro. The thread also mentions mimo-v2.6-flash; mimo-v2.6-pro-ultraspeed is a speed mode, and its latency metrics must not be folded into quality evaluations of the base Pro model.
Applicable clients, agents, or APIs: MiMo API, MiMo Desktop, and agent harnesses that can provide multimodal inputs, code execution, browser/desktop control, or other tools. The official thread also provides API, Desktop, and Hugging Face entry points.
Recommended reasoning tier and parameters: The official thread and release page do not provide a unified temperature, top-p, reasoning tier, random seed, tool timeout, or token configuration. A rerun should lock the target endpoint's defaults and record them in full; recommended parameters cannot be inferred from this source.
The main post introduces Pro and Flash as two native omnimodal models, says that the models advance capabilities through scaled reinforcement learning, and describes Pro as comparable to Claude Opus 5 and GPT-5.6 Sol on most Agent benchmarks. The post also cites a score of 46 on the Artificial Analysis Intelligence Index, calling it the highest score for an open-source model.
Visible follow-up posts add the cost framing: Pro and Flash retain the V2.5 API pricing; in comparisons that the official account considers to have “equivalent intelligence,” Pro costs about 1/20–1/60 as much as leading international models. The thread does not provide the complete comparison set, price timestamps, input/output token accounting, or the Artificial Analysis run configuration, so this is recorded only as an official promotional and positioning signal.
| Direction | Official description visible in the thread | Scope and limitations |
|---|---|---|
| 3D / Vibe World | Turn text, images, or video into playable 3D worlds; coordinate agents to build scenes, write interaction logic, and iterate on results; generate Blender objects and scenes | This demonstrates workflow capability, not a unified success rate or task sample size |
| Computer use | Use desktop tools to search, edit, and process data, then inspect results and adjust the next action | Requires a matching desktop environment, tool permissions, and observable state feedback |
| Embodied simulation | Control a Franka Panda robotic arm in simulation through visual feedback | The thread does not disclose the control frequency, simulator, success criteria, or failure samples |
| Materials research | Collaborate with Xiaomi materials researchers to propose MOF materials for capturing PFAS, search literature and patents, perform computational screening, and identify wet-lab candidates | This is an official case, not a blind-test success rate for materials discovery |
| Formal mathematics | Help formalize Li–Yorke's “Period Three Implies Chaos” main theorem in Lean 4; the official account says the integrated result exceeds 6,000 lines of code and is verified by the Lean kernel | The thread does not disclose the complete prompt, collaboration trace, or amount of human editing |
| Design and multimedia | Combine code, design, and tool use across interfaces, presentations, SVG, video, and music; claim to be comparable to Claude Opus 5 and GPT-5.6 Sol on Design Arena | No Design Arena score, task set, evaluation protocol, or original artifact collection is provided |
| Open source and reproduction | Release Pro/Flash, MiMo-V2.6-Distill-Qwen-9B, a technical report, 7K+ RL task environments, an end-to-end RL framework, and composable mini-harnesses | “Reproducible” is a release promise; actual reproduction still requires reading the repositories, pinning versions, and saving logs |
| Availability | MiMo Desktop and membership plans are available; Pro UltraSpeed is offered on Desktop and the API, with the official account claiming up to 20× higher speed; API pricing remains unchanged | 20× is a stated upper bound for the speed mode, not a general latency guarantee for base Pro |
The “blog” link in the X main post points to the Xiaomi MiMo official release page. Its complete comparison appendix gives the following MiMo-V2.6-Pro scores; the page says that higher scores are better, and marks in-house items as Xiaomi's own tests and GDPVal 2.1 (AA) as the Elo reported by Artificial Analysis.
| Category | Benchmark | Pro score | Public framing |
|---|---|---|---|
| Coding | DeepSWE v1.1 | 71.9 | Independent benchmark name; the page does not provide the complete run configuration |
| Coding | ProgramBench | 26.5 | Official comparison table |
| Coding | MiMo Code Bench | 63.2 | in-house |
| General | GDPVal 2.1 (AA) | 1673 | Artificial Analysis Elo |
| General | Toolathlon-verified | 76.9 | Official comparison table |
| General | Automation Bench v1.0.6 | 53.1 | Official comparison table |
| General | Agents' Last Exam | 31.6 | Official comparison table |
| General | Terminal Bench 4.0 | 34.9 | Official comparison table |
| General | Terminal Bench 2.1 | 89.9 | Official comparison table |
| General | OSWorld-Verified | 82.0 | Official comparison table |
| General | JobBench | 62.0 | Official comparison table |
| Visual | MiMo Visual Coding | 72.3 | in-house |
| Cyber | CyberGym | 94.0 | Official comparison table |
| Cyber | ExploitGym | 17.8 | Official comparison table |
| Cyber | ExploitBench | 47.9 | Official comparison table |
| Cyber | SEC Bench Pro | 66.3 | Official comparison table |
| Cyber | MiMo Cyber Bench | 81.7 | in-house |
The official release page also gives training-process context: Pro and Flash each completed 30 RL steps, for approximately 750k trajectories in total; Pro training cost about $2.62 million (262 × 10,000 USD), and the relative improvement in the average pass rate on training tasks was about 12%; on DeepSWE v1.1, Pro rose from 58.4 to about 72.6. This before-and-after change is a result of the official RL process and must not be interpreted as a model-swap comparison or conflated with the final scores in the table above.
Broad task coverage: The official account reports results in coding, general agents, vision, and cyber, while the thread demonstrates 3D, desktop, research, and design workflows. Pro is therefore worth prioritizing for multi-tool, long-horizon, multimodal tasks.
A clear open-source reproduction path: The thread points to a technical report, Hugging Face, and RL resources, making it possible to inspect training settings, implementation details, and model weights rather than simply accepting a leaderboard claim.
Cost and speed are product selling points: The unchanged API pricing and “up to 20×” UltraSpeed claim are useful hypotheses for cost/latency tests, but cannot be treated directly as a user's actual bill or P95 latency.
“Comparable on most Agent benchmarks” cannot, by itself, establish that the model reaches Claude Opus 5 or GPT-5.6 Sol on every agent task.
The Artificial Analysis score of 46 cannot be used to infer accuracy on an arbitrary business dataset, and the official leaderboard cannot be treated as an independent third-party evaluation.
Research cases, the Design Arena claim, and multimedia demonstrations cannot be converted into stable success rates; the thread provides no failure cases, sample sizes, or evaluator agreement data.
Results for mimo-v2.6-flash, mimo-v2.6-pro-ultraspeed, or MiMo-V2.6-Distill-Qwen-9B must not be merged into conclusions about base Pro.
Fix the mimo-v2.6-pro endpoint, client version, system prompt, tool schema, context limit, timeout, and sampling parameters; record UltraSpeed separately.
Rerun coding, Terminal, OSWorld, tool-calling, and vision tasks with the same repository, harness, and scorer, and save per-item inputs, outputs, tool traces, tokens, elapsed time, and failure types.
Build separate task sets for 3D, Blender, desktop operation, materials design, and Lean 4; save source materials, final files, compilation/simulation logs, and human acceptance results.
When comparing with other models, keep the budget, tools, context, and retry rules identical; report sample size, success rate, cost, and P50/P95 latency instead of citing only the official ranking.
Source boundary: This document uses only the specified X main post, the 6 visible Xiaomi MiMo official follow-ups, and the Xiaomi MiMo official release page directly linked from the main post/follow-ups; visible non-official quoted posts in the thread were not used as evidence.
Official-claim boundary: The figures and claims for 46/46.32, the 1/20–1/60 price ratio, 20× speed, the Design Arena comparison, case results, and benchmark scores are all Xiaomi MiMo's reporting; this document does not claim independent verification.
Method boundary: The X thread does not provide complete test prompts, task sampling, tool schemas, random seeds, scoring scripts, failure samples, or confidence intervals; although the official release page provides a comparison table and some RL training curves, they are still insufficient to reconstruct the entire evaluation.
Time boundary: Prices, available clients, model IDs, and links reflect the content visible on 2026-09-22 and may change later.
The X main post publicly states that MiMo-V2.6 includes Pro and Flash, that Pro is comparable to Claude Opus 5 and GPT-5.6 Sol on most Agent benchmarks, and that it scores 46 on the Artificial Analysis Intelligence Index.
An X follow-up publicly states that API pricing remains aligned with V2.5; the official account claims that Pro costs about 1/20–1/60 as much as leading international models at a similar intelligence level.
The follow-ups expand the visible task set to 3D worlds, Blender, Franka Panda, desktop tools, PFAS-MOF, Lean 4, front-end work, presentations, video, and music, and claim that Pro is comparable to Claude Opus 5 and GPT-5.6 Sol on Design Arena.
The thread's public open-source list includes Pro/Flash, MiMo-V2.6-Distill-Qwen-9B, a technical report, 7K+ RL task environments, an end-to-end RL framework, and a mini-harness.
Official release page linked directly from the thread: https://mimo.xiaomi.com/mimo-v2-6; API, Desktop, and Hugging Face entry points should be taken from the links currently visible on that page.
The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.
For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.
X (Xiaomi MiMo official account); Xiaomi MiMo official release page linked from the X thread · Xiaomi MiMo (@XiaomiMiMo) · Original publication date 2026-09-22 · Site edit date 2026-09-22
Open original sourceMiMo-V2.6-Pro
Download the Tabbit client to check model access