We release a sizeable monolingual Urdu corpus automatically tagged with part-of-speech tags. We extend the work of Jawaid and Bojar (2012) who use three different taggers and then apply a voting scheme to disambiguate among the different choices suggested by each tagger. We run this complex ensemble on a large monolingual corpus and release the… See the full description on the dataset page:
https://huggingface.co/datasets/bigscience-data/roots_indic-ur_urdu-monolingual-corpus.