A reproducible evaluation set covering four registers of Vietnamese
text. Used by the nom-vn project
to compare diacritic-restoration models against the public
Toshiiiii1/Vietnamese_diacritics_restoration_5th
SOTA on a register-balanced grid.
Multi-corpus measurement is the rule — single-corpus quality
numbers hide register-shift weakness. This dataset is the
multi-register grid we maintain.… See the full description on the dataset page:
https://huggingface.co/datasets/nrl-ai/vn-diacritic-eval.