Views
No views yet
Qwen/Qwen3-VL-4B-Instruct with self-supervised reinforcement learning for online proactive assistance in GridWorld environments.| Base model | Checkpoint | Speedup on GridWorld Proactive Assistance |
|---|---|---|
| Qwen/Qwen3-VL-4B-Instruct | MindZero-gw-asst-Qwen3-VL-4B-Instruct | 23.0 |
| Qwen/Qwen3-VL-8B-Instruct | MindZero-gw-asst-Qwen3-VL-8B-Instruct | 24.5 |
1@inproceedings{zhang2026mindzero,
2 title = {MindZero: Learning Online Mental Reasoning With Zero Annotations},
3 author = {Shunchi Zhang and Jin Lu and Chuanyang Jin and Yichao Zhou and Zhining Zhang and Tianmin Shu},
4 booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
5 year = {2026}
6}