Uzbek NER Gold is a token-level named entity recognition dataset for Uzbek. The dataset is distributed as a UTF-8 TSV file and uses BIO tagging for named entities.
Dataset ID: uznlp-uz/uzbek_NER
Language: Uzbek (uz)
Rows: 59,569 token rows
Columns: 5
Sentences: 4,176
Split: train
Format: UTF-8 TSV
Data file: Uzbek_NER_Gold.tsv
License: CC BY 4.0
Sentence
Sentence identifier.… See the full description on the dataset page:
https://huggingface.co/datasets/uznlp-uz/uzbek_NER.