AI Models & Platforms
Xiaomi’s New Flagship Model Leads Open-Weight Rankings With a Score of 46

Xiaomi on September 22, 2026 announced the release and open-sourcing of the MiMo-V2.6 series, led by a flagship model that benchmark owner Artificial Analysis scores at 46 on its Intelligence Index, the highest open-weights result on that leaderboard.
The series includes two natively omnimodal models: MiMo-V2.6-Pro, which Xiaomi describes as its most capable model to date, and MiMo-V2.6-Flash, which the company says strikes a balance between intelligence, efficiency, and cost. A third variant, MiMo-V2.6-Pro-UltraSpeed, delivers up to 20x faster output at the same quality for users who require extreme generation speed, according to Xiaomi.
The Artificial Analysis Benchmark Standing
Artificial Analysis’ model page for MiMo-V2.6-Pro records a 46 on Intelligence Index v4.3.2, a composite of 10 evaluations that include AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, and Humanity’s Last Exam. The benchmarker’s open-source model comparison places MiMo-V2.6-Pro first among open-weights models, ahead of Z AI’s GLM-5.3 (max) at 45 and Kimi K3 (max) at 44, and describes MiMo-V2.6-Pro and GLM-5.3 (max) as the highest-intelligence open source models.
Among proprietary systems, Artificial Analysis’ page for Grok 4.7 (xhigh) lists the SpaceXAI model at the same 46. Xiaomi’s announcement cites a 46.32 on the index and calls MiMo-V2.6-Pro “the strongest open-source model to date,” saying it surpasses Kimi K3 and Qwen3.8 Max. Artificial Analysis puts the model’s cost at $0.13 per Intelligence Index task and reports that evaluating it on the index cost $206.66. Xiaomi says the series keeps the API pricing of the V2.5 series, which it describes as pushing the intelligence-versus-cost Pareto frontier outward.
Artificial Analysis lists the model’s release date as September 21, 2026; Xiaomi’s announcement post is dated September 22, 2026.
Sparse MoE Architecture and MIT-Licensed Weights
According to the model card published by the Xiaomi MiMo Team, the flagship MiMo-V2.6-Pro-RL checkpoint is a sparse mixture-of-experts model with 1.02 trillion total parameters and 42 billion activated per token, a 1M-token context length, and text, image, video, and audio input with text output. Artificial Analysis’ specifications list the same parameter counts, context window, and modalities, and record an MIT license, which permits commercial use.
The card details a 70-layer backbone, 60 sliding-window attention layers interleaved with 10 global-attention layers, with 384 routed experts of which eight are active per token. A 681M-parameter MiMo ViT handles vision, a 308M-parameter AudioTokenizer and a 127M-parameter audio patch encoder handle audio, and a five-layer multi-token-prediction speculative decoder predicts seven subsequent tokens per forward pass.
Weights for MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL are published on Hugging Face, with ModelScope availability at release. Xiaomi said it is also open-sourcing the full technical report, the training environments, and the RL code so researchers can reproduce and verify the results.
A Livestreamed Reinforcement-Learning Run
Xiaomi said it streamed the production reinforcement-learning run live as it happened. In under six days, MiMo-V2.6-Flash and MiMo-V2.6-Pro each completed 30 RL steps over roughly 750,000 trajectories, at reported costs of about $0.85 million and $2.62 million respectively. Average pass rate on the training tasks rose by 25% and 12% in relative terms, and on DeepSWE v1.1, a held-out long-horizon software-engineering benchmark, scores rose from 48.8 to 65.68 for Flash and from 58.4 to 72.57 for Pro, Xiaomi reported.
The company said it scaled RL compute along three axes: larger batches on a fully asynchronous architecture with 1,568 samples per update, training at up to 1M context length and 3.5 to 3.7 billion tokens per step; a multi-task suite spanning coding, general agents, visual, and cyber tasks mixed across several harnesses; and more grader compute, using relative comparison within each group to produce more precise reward signals.
The model card specifies fully asynchronous Group Relative Policy Optimization at 1,568 prompts with 16 rollouts each per step, Groupwise Reward Synthesis to build task-specific rubrics offline from contrasting rollouts, and Groupwise Advantage Redistribution to rank passing trajectories online. A multi-prefix multi-teacher on-policy distillation stage called MOPD2 follows the mixed RL run, and Xiaomi describes reward-hacking defenses spanning reward design, adversarial evaluation, anomaly detection, and cross-checking between verifiers.
Pricing, Availability, and Demonstrated Uses
In Xiaomi’s own benchmark tables, MiMo-V2.6-Pro scores 71.9 on DeepSWE v1.1 against 67.9 for Flash and 19.0 for MiMo-V2.5-Pro, trailing DeepSeek V4.1 Flash (74.2), Claude Opus 5 (74.0), and GPT 6 Astra (74.0). Xiaomi lists Pro at 1,673 on GDPVal 2.1, an Elo rating reported by Artificial Analysis, behind Claude Fable 5.1 (1,735) and Claude Opus 5 (1,708). Other reported Pro results include 76.9 on Toolathlon-verified, 53.1 on Automation Bench v1.0.6, 34.9 on Terminal Bench 4.0, and 94.0 on CyberGym, where Flash posts 95.1 against GLM 5.3’s 84.5.
Both models are available in AI Studio, MiMo Code, and MiMo Desktop, as well as through the MiMo API Platform and OpenRouter. MiMo Desktop leaves early access with its first official release, with Pro and Flash built in, and Xiaomi said the UltraSpeed early-access program will run for one more week, with existing users asked to switch to the new model names.
API pricing is unchanged from the V2.5 series: per million tokens, MiMo-V2.6-Pro costs $0.0036 for cache-hit input, $0.435 for cache-miss input, and $0.87 for output; MiMo-V2.6-Flash costs $0.0028, $0.14, and $0.28; and MiMo-V2.6-Pro-UltraSpeed costs $0.036, $4.35, and $8.70 in the same three categories. Artificial Analysis measured 129.7 output tokens per second and a 2.17-second time to first token through Xiaomi’s API.
Xiaomi’s demonstrations cover multi-agent game development, Blender 3D modeling, closed-loop control of a Franka Panda robotic arm from multi-view camera feeds, frontend and presentation design, end-to-end video creation, and music composition, including an orchestral piece called Night Road that the model scored and converted to MIDI on its own. In two research cases reported by the company, MiMo-V2.6-Pro designed metal-organic framework candidates for adsorbing PFAS compounds with Xiaomi’s materials experts, and it produced a Lean 4 formalization of Li and Yorke’s “Period Three Implies Chaos” theorem exceeding 6,000 lines of code, verified by Lean’s kernel with no unfinished proof placeholders.












