A Vietnamese–Khmer translation ranking dataset for preference learning, reward modeling, and machine translation evaluation.
Vi-Km-TRank is a translation ranking dataset designed for training and evaluating preference-based machine translation models.
Unlike conventional parallel corpora that provide only a single reference translation, each example in Vi-Km-TRank contains multiple candidate Khmer translations together with a human-annotated… See the full description on the dataset page:
https://huggingface.co/datasets/Huyisbeee/Vi-Km-TRank.