This dataset hosts the Megatron-LM/MegaDLMs preprocessed SimpleStories dataset using an 8k vocab BPE Tokenizer.
Preprocessed Megatron dataset files (e.g., .bin / .idx) ready for Megatron-LM/MegaDLMs training
BPE Tokenizer config files used to create the dataset
Fork with extra preprocessing utils for SimpleStories:
https://github.com/triloy8/MegaDLMs
Original MegaDLMs repo:… See the full description on the dataset page:
https://huggingface.co/datasets/trixyL/simplestories-8k-megatron.