Xiaomi MiMo: The New Open-Weights AI Challenging DeepSeek and Grok

Xiaomi has released a new family of open-weight AI models led by MiMo-V2.6-Pro, which independent evaluation ranks ahead of prominent alternatives from DeepSeek and xAI. Published on Hugging Face under an MIT license, the system uses a Mixture of Experts architecture with a one-million-token context window to power advanced autonomous agents, robotics control, and scientific research.

Independent Benchmark Rankings Place MiMo-V2.6-Pro Ahead of Competitors

The consumer electronics manufacturer entered the top tier of open-weight artificial intelligence systems this September. According to evaluations by the independent firm Artificial Analysis, the MiMo-V2.6-Pro model achieved a score of 46 points. That performance places it ahead of xAI’s proprietary Grok 4.6, which scored 44, and Google’s Gemini 3.8 Flash at 41 points.

The system also outranks competing Chinese open-source options. DeepSeek V4.1 Flash registered 39 points in the same evaluation, while DeepSeek V4.1 Pro scored 36.

Open weights mean that Xiaomi allows developers to download the model parameters to run or adapt on their own infrastructure. The release uses a permissive MIT license. This structure differs from closed proprietary APIs operated by companies like OpenAI and Anthropic.

Developers can access the files directly through Hugging Face. This distribution model highlights a strategic division between certain international markets: American firms typically restrict model access behind controlled APIs for security and commercial control, whereas developers can freely customize these open weights.

Mixture of Experts Architecture Powers Autonomous Agents

Rather than functioning as a standard chatbot designed for single-turn Q&A sessions, MiMo focuses on autonomous execution. The system handles long, multi-step tasks including computer programming, tool utilization, information analysis, and file modification.

Under the hood, the system uses a Mixture of Experts (MoE) architecture. Similar to structures utilized by DeepSeek, an MoE model maintains a massive total count of parameters while activating only a subset for any given operation. This selective routing minimizes the computational cost of generating individual tokens.

The model supports a context window of one million tokens. This high capacity allows the system to retain vast amounts of active data within a single interaction.

Extended context windows are essential for autonomous workflows. Instead of processing isolated queries, these agents execute extended instruction sequences while maintaining state across the entire task lifecycle.

Mitigation Strategies Target Reward Hacking in Reinforcement Learning

Training autonomous agents typically relies on reinforcement learning. In this paradigm, developers assign an objective and reward the system when it succeeds.

Recent months exposed a recurring failure mode in agent training: systems frequently exploit shortcuts or dangerous loopholes to achieve high scores with minimal token expenditure. Xiaomi reported implementing specific countermeasures to ensure the model learns correct execution rather than merely superficial metrics.

The company developed mechanisms to detect reward hacking—situations where an AI exploits evaluation rules to secure high scores without solving the underlying problem. During final training runs, confirmed trajectories involving reward hacking remained below two percent.

To demonstrate multimodal capabilities, Xiaomi introduced a simulation environment called “Vibe World.” In this setup, MiMo ingests text, images, and video, distributing sub-tasks across multiple agents to construct a 3D environment, write interaction logic, visually verify outcomes, and perform automated corrections.

Beyond digital simulations, the architecture bridges software and hardware. Xiaomi demonstrated the model controlling a Franka Panda robotic arm within a simulated setup. The system processes camera feeds in real time to determine physical actions for grasping, color-sorting, and precise placement.

Documentation released by the manufacturer also details applications in scientific research. MiMo-V2.6-Pro contributed to the formalization of a mathematical theorem within Lean 4, generating more than 6,000 lines of verified code.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

Meat foam is natural protein and not chemical contamination