Big thanks to Ai2 for releasing the original PixMo-Cap dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images.
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models.
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model… See the full description on the dataset page:
https://huggingface.co/datasets/jonathanli/pixmo-cap-images.