Background of the original dataset: Localized Narratives provides a new form of multimodal image annotation that connects vision and language, collected by recording people's free narration of what they see in the pictures.
Extracted quantity: 5000 image-caption pairs are extracted from this dataset.