This is 30k randomly sampled image-captioned pairs from the COCO 2014 val split. This is useful for image generation benchmarks (FID, CLIPScore, etc.).
Refer to the gist to know how the dataset was created:
https://gist.github.com/sayakpaul/0c4435a1df6eb6193f824f9198cabaa5.