This dataset is a weakly-supervised Named Entity Recognition (NER) dataset for the Dhivehi language, built from a large unlabeled sentence corpus using dictionary-based tagging and BIO post-processing.
Language: Dhivehi (ދިވެހި) + Arabic (For Dhivehi names only)
Records: 90,735 (cleaned from 97,308 original)
Total Tokens: 775,136
Total Entities: 128,764
Data Quality: 93.2% (6,573 invalid records removed)
Average Sentence Length: 8.5… See the full description on the dataset page:
https://huggingface.co/datasets/alakxender/dhivehi-ner-dataset.