Every filtering annotation we computed for the small data pool of our
DataComp-VLM paper: image quality, image–text alignment, language ID,
text-quality classifiers, multimodal perplexity, decontamination scores and more — up to 167 fields per
sample (180 distinct fields overall), for all 120,940,134 samples across 166 source datasets.
These are the raw annotations, not a filtered dataset. They are the inputs our curation pipeline… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small_annotations.