This repository contains the Flickr30k dataset converted into the Hugging Face datasets format for easier loading and integration with modern machine learning workflows. The original dataset consists of approximately 31,000 Flickr images, each paired with five human-written captions, and is widely used for image captioning, vision-language learning, and cross-modal retrieval research. This repository preserves the original data while providing a… See the full description on the dataset page:
https://huggingface.co/datasets/AminDehnavi/flickr30k.