The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo-3-1125-32B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl.
For more information on Dolma, please see our original release here.
If you would like a smaller sample of this mix… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/dolma3_mix-6T.