Sangraha is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
More information:
For detailed information on the curation and cleaning process of Sangraha, please checkout our paper on Arxiv;
Check out the scraping and cleaning pipelines used to curate Sangraha on GitHub;
For… See the full description on the dataset page:
https://huggingface.co/datasets/ai4bharat/sangraha.