Pre-training corpus for seed models in "Scalable Data Ablation Approximations for Language Models through Modular Training and Merging", to be presented at EMNLP 2024.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/claran/seed-pretrain-decon.