The YALM Pretraining Data - 4 is a mix of English, Hindi, Math and Python Code taken from various sources for the Language modeling task and development of YALM(Yet Another Language Model).
Total Samples: 128M (~256B tokens at 2048 Context)
Test Split: 2k Samples
Shuffle Seed: 101
Datasets:
EleutherAI/SmolLM2-135M-100B
Language: English
Sources: fineweb_edu, dclm_edu, cosmopedia_v2, etc..
Hindi(20% - 25.60M):… See the full description on the dataset page:
https://huggingface.co/datasets/kp7742/YALM-pretrain4-128M.