This is a "general purposes dataset", of images that are 512px to 1024px in size, if I recall correctly
(in contrast to the "2mp" and "4mp" datasets)
Also, they are "squarish", which means that, even if they are not precisely square, the image looks fine if you
hard-crop it to force square.
Additionally, images have been hand-culled to throw out anything I considered bad for AI training.
Additionally, they were AI-culled to have a "simple photographic… See the full description on the dataset page:
https://huggingface.co/datasets/opendiffusionai/cc12m-small-squarish-simple.