Dataset Card for "vocab_filtered_dataset_22B"
Dataset Summary
This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (
https://github.com/UIUCLearningLanguageLab/AOCHILDES)
We filter the train split of SlimPajama dataset (
https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page:
https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_22B.