This dataset is a self-contained, recaptioned export of the train split from
opendiffusionai/cc12m-1mp_plus-realistic.
The upstream dataset provides metadata and image URLs; this release contains the
downloaded image bytes, so training does not require fetching images from the
original URLs.
It contains 573,052 image-caption pairs packaged as aspect-ratio-bucketed TAR
shards for text-to-image training. Images are assigned to buckets targeting a… See the full description on the dataset page:
https://huggingface.co/datasets/data-archetype/cc12m-1mp-plus-realistic-bucketed-1024.