Dataset Card for ACES and Span-ACES
Dataset Summary
ACES consists of 36,476 examples covering 146 language pairs and representing challenges from 68 phenomena for evaluating machine translation metrics. We focus on translation accuracy errors and base the phenomena covered in our challenge set on the Multidimensional Quality Metrics (MQM) ontology. The phenomena range from simple perturbations at the word/character level to more complex errors based on discourse and… See the full description on the dataset page: https://huggingface.co/datasets/nikitam/ACES.