The raw candidate pool at the small scale of our DataComp-VLM
benchmark: 120,940,134 samples / ~187.5B tokens / 10.3 TB across 166 source datasets, as
WebDataset tar shards.
This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose
the filters and the mixing ratios, and create another training set. If you instead want a
ready-to-train dataset, use dcvlm-baseline-200b
(our reference SoTA DCVLM-baseline… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small.