MorphScore is a tokenizer evaluation framework, which evaluates the extent to which a tokenizer segments words along morpheme boundaries. This repository contains the datasets used to calculate MorphScore.
In total, we have datasetes for 86 languages, but only 70 languages have at least 100 items after filtering.
All datasets are derived from existing Universal Dependencies treebanks. In the table below, we link the source dataset for each language.
See the new preprint… See the full description on the dataset page:
https://huggingface.co/datasets/catherinearnett/morphscore.