Overview.
This dataset is used for supervised fine-tuning (SFT) in training Chain-of-Focus: Adaptive Visual Search and Zooming for Multimodal Reasoning via RL,
Details.
This dataset includes 5.4k reasoning samples: 2.4k involve zoom-in behavior, and 3k are text-only reasoning cases. All samples were generated using our Visual Agent and span a variety of image resolutions.
Training Code: The SFT code can be found at
https://github.com/xtong-zhang/Chain-of-Focus
Project page:… See the full description on the dataset page:
https://huggingface.co/datasets/xintongzhang/CoF-SFT-Data-5.4k.