ROCS is a text-to-image retrieval benchmark for small, rare objects in cluttered
scenes. Cluttered COCO and Flickr30K images are filtered with SAM3 segmentation to
find single-instance, visually subordinate objects, then re-captioned with Qwen3-VL to
name one such object per caption. Retrieval is therefore tested on whether the named rare
object is found, not just the dominant scene. Two configs: coco (3,248 images, 8,231
queries) and… See the full description on the dataset page:
https://huggingface.co/datasets/AbdulmalekDS/ROCS.