Uzbek POS: First UPOS tagged dataset for Part-of-Speech tagging task
This dataset is an annotated dataset for POS tagging. It contains 250 sample sentences collected from news outlets and fictional books respectively.
The dataset is presented in both Uzbek scripts i.e., Latin and Cyrillic. The annotation was done manually according to UPOS tagset.
Languages
Northern Uzbek (a.k.a Uzbek)
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/latofat/uzbekpos.