Xiaomi (小米) Releases Self-Evolving Large Model System MiMo-V2.6
Xiaomi (小米) MiMo team member Luo Fuli announced the release of MiMo-V2.6, described as one of the largest single-run reinforcement learning training efforts by an open-source model team to date. The team utilized MixRL for verifiable tasks and separate training for complex, long-cycle tasks, merging capabilities via MOPD. To support agent reinforcement learning research, the team also released a Qwen model distilled from MiMo trajectories, 7,000 diverse environments, and a complete reinforcement learning training framework.
Summaries are written by AI from the original article. Not investment advice.