This release contains two COCO 2014-derived splits for visual grounding with
reference-style horizontal rows of repeated non-person objects.
coco2014_candidates_4to8_unfiltered/: 15,577 candidate images. Each has
one selected non-person COCO category with 4-8 annotated instances.
strict_ready_131/: 131 visually screened samples with complete grounding
annotations. Every item passed whole-image row screening and individual
crop validation.… See the full description on the dataset page:
https://huggingface.co/datasets/cuteLyra/RefStuff.