An Urdu text corpus to enable research and applications for the Urdu language. We believe Maḵẖzan is the best Urdu dataset to start work for Urdu NLP.
This dataset currently comprises 6.26 million words of Urdu text. We have selected source text that we believe to have gone through strong editorial standards, to preserve linguistic integrity. The text is then syntactically marked up, so that headings, paragraphs, and lists can be identified. Metadata is added to each… See the full description on the dataset page:
https://huggingface.co/datasets/ReySajju742/makhzan-urdu.