This dataset, ALIF_Urdu_Corpus, is part of the ALIF الف project by Orature AI. It was curated for pretraining Urdu language models. It serves as a preview to our entire 33GB Dataset.
Purpose of the Dataset:
(preview) To serve as a large-scale, diverse, and high-quality foundation for pretraining generative language models for Urdu.
The data is in Urdu.
(For Pretraining Corpus - ALIF-Urdu-Corpus):… See the full description on the dataset page:
https://huggingface.co/datasets/sarasarahuss/ALIF_Urdu_Corpus_AUC.