We work with the German-English portion of the QT21 dataset (Specia et al. 2017), which originally contains a total of 43,000 examples of machine translation human post-edits. The machine translation output over which post-editing is performed to create this dataset is an implementation of the attentional encoder-decoder architecture and uses byte-pair… See the full description on the dataset page:
https://huggingface.co/datasets/jvamvas/peer_qt21-de-en-pe.