Overview.
The SFT data for training Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning,
The queries require fine-grained visual analysis in both images (e.g., infographics, visually-rich scenes, etc) and videos.
Details.
The data contains 8,000+ reasoning trajectories, including :
2,000+ textual reasoning trajectories, rejection sampled from the base model Qwen2.5-VL-Instruct. These data aims to preserve textual reasoning ability on easier VL… See the full description on the dataset page:
https://huggingface.co/datasets/TIGER-Lab/PixelReasoner-SFT-Data.