DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool.
The instruction-heavy counterpart (DCVLM-baseline) is available as
dcvlm-baseline-200b, along with… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.