A cleaned, morphologically-annotated dictionary of content words built from
Wiktionary, intended as raw material for designing morphology-aware tokenizer
benchmarks. There is one record per word (parts of speech merged): each record
carries the word's POS list, frequency signals, a morphological-structure label, the
full inflectional paradigm, definitions with examples, and translations. The schema is
flat (all fields at… See the full description on the dataset page:
https://huggingface.co/datasets/yuanxin112/wiktionary-morph.