Tabbit
ResourcesBlogModels
Tabbit LogoTabbit

Tabbit — The AI Browser that Works for You

Topics

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

Popular Guides

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

Events

  • Tabbit Skill Competition
  • KPOP SBTI Fandom Personality Test
  • Tabbit Campus Creator Program
  • fifi's Picks: AI Skills for Research Papers
  • User Survey

About

  • Tabbit Blog
  • Press & Media
English
简体中文English
Reviews and evidence

MiMo-V2.6-Pro · Media / benchmark · Vendor report

Xiaomi MiMo Official Release: MiMo-V2.6-Pro Benchmark Signals and Native Omnimodal Positioning

Xiaomi positions MiMo-V2.6-Pro as a native omnimodal open-source model for agents, coding, vision, and computer use, and reports an Artificial Analysis Intelligence Index score of 46 along with several RL/agent results; these figures remain the vendor's own reporting and cannot replace an independent rerun under the same harness.

Media / benchmarkVendor reportEdited 2026-09-22

Test conditions

Source-specific observation
Xiaomi positions MiMo-V2.6-Pro as a native omnimodal open-source model for agents, coding, vision, and computer use, and reports an Artificial Analysis Intelligence Index score of 46 along with several RL/agent results; these figures remain the vendor's own reporting and cannot replace an independent rerun under the same harness.
Published conditions
Deciding whether to ship to production based only on official leaderboards; rigorous comparisons requiring independent statistical significance, complete inputs, random seeds, and per-item logs; or treating case demonstrations and claims about what the model “can achieve” as stable success rates.

Key data and applicable tasks

One-sentence conclusion

Xiaomi positions MiMo-V2.6-Pro as a native omnimodal open-source model for agents, coding, vision, and computer use, and reports an Artificial Analysis Intelligence Index score of 46 along with several RL/agent results; these figures remain the vendor's own reporting and cannot replace an independent rerun under the same harness.

Suitable use cases

  • Suitable tasks: Long-horizon execution and tool use for software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development; building 3D scenes from images, video, or text; Blender modeling; visually guided robot control; desktop office work and data processing; materials research; formal proofs; and front-end, presentation, video, and music creation.

  • Unsuitable tasks: Deciding whether to ship to production based only on official leaderboards; rigorous comparisons requiring independent statistical significance, complete inputs, random seeds, and per-item logs; or treating case demonstrations and claims about what the model “can achieve” as stable success rates.

  • Applicable model versions: mimo-v2.6-pro. The official page also lists mimo-v2.6-flash; mimo-v2.6-pro-ultraspeed is a high-speed mode identifier in the API and should not be conflated with evaluation results for the base Pro model.

  • Applicable clients, agents, or APIs: Xiaomi MiMo API, MiMo Desktop, and agent harnesses that provide multimodal inputs, tool use, browser/desktop access, or code execution. The official page does not provide a complete, directly reusable tool schema.

  • Recommended reasoning tier and parameters: The official release page does not publish a consistent temperature, top-p, reasoning tier, token limit, tool timeout, or random seed. For a rerun, use the target endpoint's defaults and record each setting; do not infer recommended parameters from this page.

Test method/workflow

This is a vendor release note, not an independent rerun. Based on the experimental clues that can be checked on the official page, the following process is recommended for reproduction:

  1. Fix the API model ID mimo-v2.6-pro, client version, tool schema, context limit, timeout, sampling parameters, and harness; do not fold latency results for mimo-v2.6-pro-ultraspeed into the base model's results.

  2. Rerun tasks from SWE-bench Verified, DeepSWE v1.1, Terminal Bench 2.1, and similar benchmarks on the same code repository and scoring script, while saving per-item outputs, pass rates, tool-call counts, input/output tokens, elapsed time, and failure types.

  3. For the agent capabilities claimed by the official page, test coding, general knowledge, vision, and cyber tasks separately; report the sample size for each subset instead of reporting only that the model is comparable to another model on “most Agent Benchmarks.”

  4. For omnimodal capabilities, use reproducible tasks: generate 3D scenes from image/video/text inputs, perform Blender modeling, complete grasp-and-place actions from multi-view images, and perform desktop information retrieval, editing, and data processing. Save input materials, tool traces, final files, and human acceptance results.

  5. For research cases such as materials design and Lean 4 formalization, record retrieval sources, tool versions, prompts for every round, model edits, compilation/simulation logs, and human review. Do not turn official case descriptions into unverified success rates.

  6. When comparing with other models, use the same inputs, tools, budget, and scoring rules; report the vendor's AA score, training improvements, and costs separately, marked as “not independently reproduced.”

Original evidence and data

1. Official positioning and public model interface

  • The page says that the MiMo-V2.6 series consists of Pro and Flash, two “native omnimodal” models, emphasizing recursive self-improvement (RSI) and scaling reinforcement learning (RL) on verifiable complex tasks.

  • The official API documentation requires lowercase model names: mimo-v2.6-pro, mimo-v2.6-flash, and mimo-v2.6-pro-ultraspeed.

  • The page extends the capability scope to 3D spatial reasoning, multimodal perception, and Computer Use Agent (CUA), and describes natural-language programming as expanding into “Vibe World,” a way to build interactive worlds. These are capability-positioning claims, not scores from a unified benchmark.

2. Officially reported benchmark and RL data

Metric or settingInformation given on the official MiMo-V2.6-Pro pageEvidence boundary
Artificial Analysis Intelligence Index (AA Composite Intelligence Index)46 pointsSelf-reported on the vendor's release page; the page says it exceeds Kimi K3 and Qwen3.8 Max, but does not provide per-item inputs, configurations, or independent verification logs in the body
DeepSWE v1.1 (long-horizon software engineering, out-of-sample)58.4 → 72.6, an increase of about 14 pointsThe page attributes the change to about 6 days and 30 steps of RL training; it cannot be interpreted as a gain caused only by replacing the model
RL training steps30 stepsThe page says that Pro and Flash each completed 30 steps
Cumulative trajectoriesAbout 750,000The page describes about 750,000 trajectories for the two models combined, without a Pro/Flash breakdown
Pro RL training costAbout $2.62 million (the source figure is 262 × 10,000 USD)Vendor-reported training cost, not per-inference cost or the user's actual bill
Average pass rate on Pro training tasksAbout 12% relative improvementThe page does not provide a task list, absolute baseline, confidence interval, or per-task results
Batch size per update1,568 samplesDescription of the official training setup; it is not the inference-time context or batch configuration
Training context lengthSupports 1M contextDescription of the training setup; it cannot directly be treated as a promise that every API endpoint supports this context length
Tokens per training stepAbout 3.5–3.7BDescription of the official training setup

The same section of the official explanation also gives a Flash comparison: DeepSWE v1.1 increased from 48.8 to 65.7 (about 17 points), the average pass rate improved by about 25% relative, and training cost was about $850,000 (the source figure is 85 × 10,000 USD). These figures explain the official RL narrative and should not be treated as evaluation results for Pro.

The page also says that MiMo-V2.6-Pro is comparable to Claude Opus 5 and GPT-5.6 Sol on “most Agent Benchmarks,” but the body does not list the benchmark names, scores, sample sizes, or run configurations. This can therefore be recorded only as official positioning, not formatted as a comparable leaderboard table.

3. Omnimodal and agent cases

  • 3D open world: The official description says the model decomposes image-, video-, or text-based requirements into subtasks such as 3D scene construction, interaction-logic programming, and visual verification, continuously revises the result based on rendering output, and ultimately generates a runnable interactive world.

  • Blender: The official page says the model can build objects and scenes from text or reference images, generating assets for animation, 3D printing, and game development.

  • Embodied intelligence: The official page says that in a simulated environment, the model receives multi-view camera images and uses a visual-feedback loop to control a Franka Panda robotic arm, completing grasping, color matching, and precise placement.

  • Computer Use Agent: The official page says the model can understand complex GUIs, use office and productivity tools for information retrieval, editing, and data processing, and inspect results, troubleshoot issues, and adjust subsequent actions based on visual feedback.

  • Research cases: The official page presents MOF material design and “dry-lab” screening for PFAS adsorption targets, as well as a Lean 4 formalization of the Li–Yorke theorem “Period Three Implies Chaos.” For the latter, the page says the final output contained more than 6,000 lines of Lean source code, was verified by the Lean kernel, and had no unproven placeholders. This is an official case description, not a public blind test.

  • Content creation: The official page says that on Design Arena, Pro is at a comparable level to Claude Opus 5 and GPT-5.6 Sol, and shows examples involving front-end work, presentations, video, SVG, and music. The body does not provide Design Arena scores or a complete evaluation protocol.

4. Open-source resources and auxiliary benchmarks

The official page provides open-source information for the weights, technical report, RL training environment, and code, and says that these include 7k+ high-quality RL task environments covering four types of agent tasks: software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development.

The page also reports improvements on all 11 benchmarks starting from the MiMo-V2.6-Distill-Qwen-9B SFT baseline, including: SWE-bench Verified 61.1 → 66.2, MiMo Cyber Bench 31.3 → 47.0, Terminal Bench 2.1 37.1 → 52.8, and MiMo Visual Coding 64.0 → 72.4. These are auxiliary results from the Distill-Qwen-9B training resources and must not be treated as direct scores for MiMo-V2.6-Pro.

Scope and limitations

  • Vendor self-reporting: The AA score of 46, agent-benchmark comparisons, DeepSWE improvement, training cost, trajectory count, training throughput, Design Arena positioning, and case results all come from the Xiaomi MiMo release page. This document does not present them as independently verified.

  • Missing reproduction requirements: The body does not publish per-question inputs, the complete system prompt, tool schemas, random seeds, scoring scripts, failure samples, token/latency logs, or confidence intervals. Therefore, win rates, cost efficiency, and statistical significance cannot be calculated from the page.

  • Limits on attributing training gains: The 58.4→72.6 DeepSWE change is a before-and-after result in the official RL process and may be affected by data, environments, graders, the harness, and training strategy at the same time; it cannot be attributed to a single model version.

  • Context and API limits: 1M context is a training description, not evidence that every API, Desktop, or UltraSpeed endpoint exposes 1M; the actual endpoint should be checked against the current API documentation and validated with an overlong input.

  • Limits on price conclusions: The page says that at an equivalent intelligence level, Pro costs roughly 1/20 to 1/60 of overseas models, but the body does not provide the complete price table, model set, token accounting, or time window used for the comparison. The ratio can only be treated as official promotional language.

  • Cases do not equal stable capability: The 3D, robotics, research, formal-proof, video, and music cases lack a unified sample and failure rate. Before production use, retain artifacts, logs, and human acceptance results.

  • Version boundaries: Do not describe the maximum 20× inference speed of mimo-v2.6-pro-ultraspeed as a quality improvement of the base Pro model, and do not merge Flash or Distill-Qwen-9B results into Pro's score.

Source excerpts or observations

  • The page explicitly states that MiMo-V2.6-Pro scored 46 on the AA Composite Intelligence Index and claims to exceed Kimi K3 and Qwen3.8 Max. This sentence is the page's most direct public ranking evidence.

  • The page gives Pro's DeepSWE v1.1 result as 58.4 → 72.6, alongside context including about 6 days, 30 RL steps, and about $2.62 million (the source figure is 262 × 10,000 USD) in training cost. A review should therefore record the “model result” and the “training-process result” separately.

  • The page uses “Vibe World” to summarize the expansion from code generation to interactive-world construction and lists 3D, Blender, robotic-arm, and CUA cases. These observations support the product positioning of “native omnimodal + agent workflow,” but do not provide a unified pass rate.

  • Official open-source entry point: https://huggingface.co/collections/XiaomiMiMo/mimo-v26. The benchmark figures in this document remain based on the release-page body; potentially updated model-card content at that entry point has not been mixed into this source.

What this supports

  • Long-horizon execution and tool use for software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development; building 3D scenes from images, video。

What this does not support

  • Deciding whether to ship to production based only on official leaderboards; rigorous comparisons requiring independent statistical significance, complete inputs, random seeds, and per-item logs; or treating case demonstrations and claims about what the model “can achieve” as stable success rates.

Method, limits, and reproduction

The figures, task set, reasoning tier, and client conditions apply only to the listed source and collection snapshot. Different versions, harnesses, or providers must not be compared directly; undisclosed parameters remain unknown.

For a reproduction, fix the model version, provider or client, reasoning tier, tools, task-set version, sample count, and collection date, and record failures, retries, and human corrections. Full steps are in the source notes below.

Original source

Xiaomi MiMo Documentation · Xiaomi MiMo · Original publication date 2026-09-22 · Site edit date 2026-09-22

Open original source

MiMo-V2.6-Pro

Compare MiMo-V2.6-Pro in Tabbit

Download the Tabbit client to check model access

Read the full analysis

Full review · English

MiMo-V2.6-Pro Review: The Smartest Open Model Makes You Wait

A public-evidence review of MiMo-V2.6-Pro: what it does well, where it bites, real user reports, and a workload verdict on Xiaomi's open flagship.

Pricing · English

MiMo-V2.6-Pro Pricing: Official Rate Card, Cache Levers, and Cost per Task

A practical decision guide to MiMo-V2.6-Pro pricing: official API rates, prompt cache economics, reasoning token overhead, UltraSpeed mode, and worked task budgets.

Alternatives · English

MiMo-V2.6-Pro Alternatives: Choose by Task and Budget

Compare five MiMo-V2.6-Pro alternatives by completed-task cost, agentic reliability, open weights, and deployment fit, with prices checked on September 22, 2026.

Comparison · English

MiMo-V2.6-Pro vs MiMo-V2.6-Flash: Which Xiaomi MoE Model Fits Your Workload?

A head-to-head comparison of MiMo-V2.6-Pro and Flash: 1.02T vs 309B MoE architecture, 3.1x pricing delta, reasoning token overhead, agent benchmarks, and decision matrix.

Related reviews

MiMo-V2.6-Pro Official Technical Report: Architecture, Scaled RL, and Evaluation ConditionsThe official report defines MiMo-V2.6-Pro as a native multimodal sparse MoE with 1.02T total parameters and approximately 42B active parameters, and reports strong agent benchmark results from large-batch, multi-environment RL training with multiple harnesses and groupwise graders; however, most figures in the tables are vendor-reported。MiMo-V2.6-Pro Official X Release Thread: Task Positioning, Public Benchmarks, and Open-Source Entry PointsThis official thread positions MiMo-V2.6-Pro as an openly built native omnimodal agent model focused on coding, computer use, 3D, design, research, and tool workflows; its rankings and examples help identify promising task directions, but they remain vendor-reported and cannot replace an independent rerun under the same harness.Arena Code Arena: MiMo-V2.6-Pro WebDev AutoEval RecordArena's Code Arena | WebDev overall leaderboard includes mimo-v2.6-pro with an AutoEval score of 1628 (+18/-18), but it does not publish a vote count or rank, so this only shows that it was included in the WebDev automated evaluation leaderboard; 1628 must not be treated as a blind-test ranking.Reddit Field Test: MiMo-V2.6-Pro Gets Stuck in a grep Infinite Loop During a UI/UX Terminal Modification TaskIn a real website UI/UX modification task, the poster said that both MiMo-V2.6-Pro and MiMo-V2.6-Flash triggered a grep-related infinite loop in the terminal and failed to fix the code after approximately 30 minutes; when the same modification was assigned to DeepSeek V4.1 Flash, it continued as expected and output a complete terminal work history.Xiaomi MiMo-V2.6-Pro Omnimodal Input and Visual Task Workflowmimo-v2.6-pro can read publicly accessible URLs or properly formatted Base64 images, videos, and audio through the OpenAI Chat Completions API, but it cannot directly upload local files, and the combined media and text tokens remain subject to the 1M context limit.Xiaomi MiMo-V2.6-Pro Official API Integration and Reasoning ConfigurationThis official configuration can be used to connect to mimo-v2.6-pro through the OpenAI-compatible protocol, with deep thinking, streaming output, and multi-turn tool calls enabled as needed.Hugging Face Official MiMo-V2.6-Pro-RL Local Deployment and Chat Template ConfigurationThe official model card provides SGLang and vLLM service commands for MiMo-V2.6-Pro-RL and defines chat-template behavior for text, image, video, audio, thinking, and tool calls in the repository tokenizer configuration; local deployment must use the checkpoint name and must not treat it as the same model identifier as the hosted API's mimo-v2.6-pro.Xiaomi MiMo-V2.6-Pro Official Function Calling and Multi-Turn Agent WorkflowFor mimo-v2.6-pro, the official workflow is “the model returns a complete assistant message, including reasoning_content and tool_calls → the client executes the tools → appends the role: tool results → requests the model again,” repeating until the current turn produces no more tool calls.