This dataset contains synthetic captions, embeddings, and metadata for DFNDR-12M.
The metadata has been generated using pretrained image-text models on DFN-12M, a uniformly sampled subset of 12.8M samples from DFN-2B.
For details on how to use the metadata, please visit our ml-mobileclip repository.
For code to generate multi-modal reinforced datasets at large scale see ml-mobileclip-dr repository.
A BFloat16 version of this dataset is available at… See the full description on the dataset page:
https://huggingface.co/datasets/apple/DFNDR-12M.