-
Evidence-Preserving View: keep only the visual evidence needed to answer, reduce distractions.
→ enforce consistency between predictions from the original image and the preserved view.
-
Evidence-Ablated View: remove the key evidence so the image no longer supports the answer.
→ enforce separation so the model cannot rely on shortcuts.
1@article{zhang2025bips,
2 title={See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning},
3 author={Zhang, Shuoshuo and Zhang, Yizhen and Fu, Jingjing and Song, Lei and Bian, Jiang and Yang, Yujiu and Wang, Rui},
4 journal={arXiv preprint arXiv:2512.22120},
5 year={2025}
6}