The YALM Pretraining Data - 3 is a mix of Math, Python Code and Multilingual Data in English, Hindi and Gujarati taken from various sources for the Language modeling task and development of YALM(Yet Another Language Model).
Total Samples: ~122M
Test Split: 22k Samples
Note: This Dataset is not shuffled but concatenated due to resource constraints. Please consider shuffling before using it.
Datasets:
HuggingFaceFW/fineweb - Language: English | Subset:… See the full description on the dataset page:
https://huggingface.co/datasets/kp7742/YALM-pretrain3-122M.