DataConcept-128M is a multimodal pretraining dataset comprising 128M web-crawled image-text pairs, derived from DataComp-CLIP annotated with fine-grained details about their concept composition. This dataset is designed to enable Concept-Aware Batch Sampling (CABS), a flexible batch sampling framework that constructs batches on-the-fly based on⦠See the full description on the dataset page:
https://huggingface.co/datasets/bethgelab/dataconcept_128M.