This is the SophiaVL-R1-Thinking-156k dataset for training Thinking Reward Model of SophiaVL-R1 (SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward).
The data is constructed in sharegpt format. text_only_part.json is text-only data. multimodal_part.json is image-text data. Images can be found in images.