This dataset is in WikiAnn format. The dataset is assembled from named entities parsed from Wikipedia, Wiktionary and Dbpedia. For some words, new case forms have been created using Apertium-uig. Some locations have been translated using the Google Translate API.
The dataset is divided into two parts: train and extra. Train has full sentences, extra has only named entities.
Tags: O (0), B-PER (1), I-PER (2), B-ORG (3), I-ORG (4), B-LOC (5)… See the full description on the dataset page: https://huggingface.co/datasets/codemurt/uyghur_ner_dataset.