This dataset contains a 12M subset of DataComp-1B-BestPool.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Image-text models trained on DataComp-12M are significantly better than on CC-12M/YFCC-15M as well as DataComp-Small/Medium.
DataComp-12M was introduced in MobileCLIP paper and along with the reinforced dataset DataCompDR-12M.
The UIDs… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations/DataComp-12M.