A word-level morphological segmentation dataset for Nepali, mapping surface
words to their root and affix components (e.g. अँकाइनु → अँका + इनु).
Dataset Details
Dataset Description
This dataset was built by scraping the 10th edition of the Nepali
dictionary, parsing the scraped entries into structured word/root pairs,
and then applying a heuristic rule engine to extend root-affix coverage
to surface forms not explicitly… See the full description on the dataset page: https://huggingface.co/datasets/W4ashabii/root_affix_dictionary.