The YALM Pretraining Data - 6 is a mix of English, Hindi, Math and Python Code taken from various sources for the Language modeling task and development of YALM(Yet Another Language Model).
Total Samples: 62M (~42B tokens with sample packing at 2048 Context)
Test Split: 10k Samples
Shuffle Seed: 101
Datasets:
zicsx/mC4-Hindi-Cleaned… See the full description on the dataset page:
https://huggingface.co/datasets/kp7742/YALM-pretrain6-62M.