The code used to retrieve and create this dataset is almost identical to the one that you can find here wmt14.
I only added the possibility to retrieve the "es-en" translation pairs from the newstest2013. This pair works only for the train and validation splits.
Pay attention: some es-en pair sentences on the validation set contain the backslash followed by a double quote character (").
Thanks to the Huggingface team for all the work they have done!