This repo contains SynthCI-30M which is the dataset proposed in "SynthCLIP: Are We Ready For a Fully Synthetic CLIP Training?". The dataset contains 30M synthetic text-image pairs covering a wide range of concepts.
We present SynthCLIP, a novel framework for training CLIP models with entirely synthetic text-image pairs, significantly departing from previous methods relying on real… See the full description on the dataset page:
https://huggingface.co/datasets/hammh0a/SynthCLIP.