Note: this is an identical copy of
https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Open weights, closed datasets… See the full description on the dataset page:
https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.