This dataset includes 1350 images of objects of varied counts, 150 images for each count between 2 and 10. It consists of images automatically sourced from multiple sources, including the COCO Dataset, Conceptual 12M, YFCC100M, and SBU Captions Dataset. In addition, due to the fact that images with large counts (i.e., nine and ten) are scarce, to compose 150 images for each count, we also manually collect some images from the Internet. Each text image pair is… See the full description on the dataset page:
https://huggingface.co/datasets/ruisu516/DiverseCount.