This dataset is a cleaned subset of the Flickr8k corpus (sourced from wds_flickr8k) for multimodal large language model pretraining. It has been processed through resizing, deduplication, and filtering to produce training-ready image-text data. Due to its relatively small size, the dataset is intended for simple validation of multimodal pretraining experiments rather than large-scale training.