Training data created for constructing quality estimation and automatic post-edition models (Ive et al. 2020). The data consists of tuples (source sentence, machine translation output, manual post-edition, independent reference translation) for three European language pairs. The data cover the following domains: online dispute resolution, procurement and justice. Number of tuples per… See the full description on the dataset page:
https://huggingface.co/datasets/FrancophonIA/multilingual_legal_corpus.