The raw candidate pool at the medium scale of our DataComp-VLM
benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as
WebDataset tar shards — ≈4× the small pool.
This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose
the filters and the mixing ratios, and create another training set. If you instead want a
ready-to-train dataset, use dcvlm-baseline-200b
(our reference SoTA… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.