William Kalikman*, Šimon Sukup*, Michal Tešnar, Vilém Zouhar (*equal contribution)
ETH Zurich
This dataset contains the (original, augmented) text pairs produced by the two ATO variants. The method, code, and configs can be seen in the repository:
https://github.com/BreakingMT/ATO
original_sentence — source English sentence (FLORES200 or WMT22/23/24 news test set)… See the full description on the dataset page:
https://huggingface.co/datasets/wskal/ATO-datasets.