A curated image captioning dataset containing 1,000 image and caption pairs. Each sample consists of an image paired with a detailed natural language description, making it suitable for training and evaluating image captioning and vision-language models.
1,000 image-caption pairs
Detailed natural language captions
High-quality descriptive annotations
Suitable for image captioning and multimodal learning