Precomputed image-tower features for 10,968,539 CC12M images (all 2,176
shards of
pixparse/cc12m-wds)
from ten independent teacher extractions — eight CLIP variants across
three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus
one derived consensus target.
About 110 million feature vectors, roughly 130 GPU-hours of extraction,
so that a student can be distilled against any of these… See the full description on the dataset page:
https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.