ImageNet-Think 250K is a large-scale synthetic multimodal reasoning dataset containing of 250,000 images sampled from ImageNet-21K dataset. For each image, we provide a prompt and two different step-by-step reasoning tokens and outputs (answers), enabling evaluation and training for Vision Language Models on reasoning tasks. This dataset is primarily designed for research on multimodal summarization.
Before downloading… See the full description on the dataset page:
https://huggingface.co/datasets/krishnateja95/ImageNet-Think.