Views
No views yet
Qwen/Qwen3-VL-4B-Instruct with self-supervised reinforcement learning for online Theory-of-Mind reasoning in gridworld environments.| Base model | Checkpoint | Gridworld-QA |
|---|---|---|
| Qwen/Qwen3-VL-4B-Instruct | MindZero-gw-tom-Qwen3-VL-4B-Instruct | 95.0 |
| Qwen/Qwen3-VL-8B-Instruct | MindZero-gw-tom-Qwen3-VL-8B-Instruct | 92.3 |
1@inproceedings{zhang2026mindzero,
2 title = {MindZero: Learning Online Mental Reasoning With Zero Annotations},
3 author = {Shunchi Zhang and Jin Lu and Chuanyang Jin and Yichao Zhou and Zhining Zhang and Tianmin Shu},
4 booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
5 year = {2026}
6}