Gemstones Training Dataset - Parallel workers sharded version
This data is a reprocessed version of the first 1B rows of the Dolma v1.7 dataset (
https://huggingface.co/datasets/allenai/dolma).
The data is encoded using the Pythia tokenizer:
https://huggingface.co/EleutherAI/pythia-160m
Disclaimer: this is an approximation of the dataset used to train the Gemstones model suite.
Due to the randomized and sharded nature of the distributed training code, the only way to perfectly
reproduce the… See the full description on the dataset page:
https://huggingface.co/datasets/tomg-group-umd/gemstones_data_order_parallel.