This dataset pairs UniMorph inflectional features with UniSegments segmentations. For languages without UniSegments coverage, segmentation defaults to the unsegmented word form itself.
This resource is a necessary component for evaluating Tokenizer Morphological Plausibility, as introduced in Tokenizer Morphological Plausibility (
https://arxiv.org/abs/2601.18536). The data generation process follows the implementation provided in the official… See the full description on the dataset page:
https://huggingface.co/datasets/SHENJJ1017/morph_features.