This is the training data of our paper: Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning, including:
SFT data for two thinking modes: text-based thinking and visually-grounded thinking (with bounding boxes)
RL data containing questions for geometric, object counting, OCR, Chart, Grounding, Science...
Please refer to our GitHub repo:… See the full description on the dataset page:
https://huggingface.co/datasets/ZejunLi/MoVT-Train.