Region-level grounding for Conceptual Captions 3M: bounding boxes, the noun
phrase each box grounds, and the span of the caption that phrase came from, for
3,016,640 of CC3M's 3,318,333 rows.
No images here. This is metadata only, joinable onto a CC3M copy you already
have. That is the point of it: the grounding is 354 MB, the pixels are 125 GB.
annotations-0000..0482.parquet
3,016,640
199 MB… See the full description on the dataset page:
https://huggingface.co/datasets/freek23/cc3m-grounded-annotations.