Project page
STRAP is a large-scale source of structured region supervision for
vision-language models. It contains annotations for 2,000,000 web images:
1M from Conceptual Captions 3M (CC3M)
and 1M from DataComp-1B.
Across both subsets, STRAP describes 10.35 million objects.
Each object connects a normalized bounding box to a basic-level label, naming
hierarchy, short description, attributes, visible parts, and categories the… See the full description on the dataset page:
https://huggingface.co/datasets/vrg-prague/STRAP.