This dataset is continuously updated and contains a compilation of human translation quality assessment from past WMT campaigns.
Specifically, this dataset merges all annotation protocols (DA, MQM, ESA) on a semi-unified scale (0 to 100).
The current version of the dataset includes human scores up to WMT 2025 (inclusive) and has been created with the following script:
import subset2evaluate # version 1.0.20
import json
import statistics
data =… See the full description on the dataset page:
https://huggingface.co/datasets/zouhar/wmt-human-all.