Welcome to Xiaomi MiMo-VL-Miloco — the first open-source multimodal model built to actually understand what’s happening at home!
We use a carefully tuned two-stage pipeline to nail home-scene skills without sacrificing general abilities.
This stage focuses on boosting the model’s core capabilities in home scenarios. Even with a limited training set, we strike a good balance between sample-efficient learning and fast inference:
Building on fine-tuning, this stage introduces GRPO-based reinforcement learning to enhance the model’s overall performance:
In short: Xiaomi MiMo-VL-Miloco is your friendly, sharp-eyed model roommate—great at recognizing what’s going on around the house, and still ready for the wider world.
In household scene understanding, we prioritize video and image perception alongside the model’s reasoning ability.
1@misc{xiaomimimovlmiloco,
2 author = {Jiaze Li, Yuxun Qu, Jingyang Chen, Shijie Xu, Zhenru Lin, Junyou Zhu, Boshen Xu, Wenhui Tan, Pei Fu, JianZhong Ju, Zhenbo Luo, Jian Luan},
3 title = {Xiaomi MiMo-VL-Miloco},
4 year = {2025},
5 howpublished = {\url{https://github.com/XiaoMi/xiaomi-mimo-vl-miloco}},
6}
Please contact us at
milm-plus@xiaomi.com or open an issue if you have any questions.