Captions and grounded bounding boxes for the ImageNet-1k train split.
Images are not included. Each annotation is keyed by its ImageNet train
image id (e.g. n13133613_29204); pair the annotations with an ImageNet-1k
copy or with
visual-layer/imagenet-1k-vl-enriched.
1,273 JSON shards covering 1,277,474 captioned samples, of which
801,472 carry grounded boxes (1,040,313 boxes in total). Boxes are
filtered at grounding… See the full description on the dataset page:
https://huggingface.co/datasets/freek23/imagenet-enriched-grounded-boxes.