Views
No views yet
| Model | V-Star | HR-Bench-4K | HR-Bench-8K | MME-RealWorld-Lite |
|---|---|---|---|---|
| Qwen3-VL-Instruct-2B | 74.9 | 70.4 | 64.6 | 47.3 |
| P2R-2B | 84.3 | 75.1 | 74.8 | 51.3 |
| Δ | +9.4 | +4.7 | +10.2 | +4.0 |
1from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
2
3model = Qwen3VLForConditionalGeneration.from_pretrained("hongxingli/P2R-2B")
4processor = AutoProcessor.from_pretrained("hongxingli/P2R-2B")1@misc{li2026perceivetoreasondecouplingperceptionreasoning,
2 title={Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning},
3 author={Hongxing Li and Xiufeng Huang and Dingming Li and Wenjing Jiang and Zixuan Wang and Haolei Xu and Hanrong Zhang and Haiwen Hong and Longtao Huang and Hui Xue and Weiming Lu and Jun Xiao and Yueting Zhuang and Yongliang Shen},
4 year={2026},
5 eprint={2607.01191},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV},
8 url={https://arxiv.org/abs/2607.01191},
9}