The Makhzan Urdu Corpus is a high-quality, expert-curated collection of classical and modern Urdu texts. This dataset has been extracted and cleaned from the Makhzan GitHub repository and structured into the Hugging Face datasets library format.
The corpus contains over 6.26 million words of Urdu text from various literary and scholarly publications. Each document is tokenized into paragraphs and tagged with semantic structure… See the full description on the dataset page:
https://huggingface.co/datasets/m-aliabbas1/makhzan-urdu-corpus.