The WMT14 injected synthetic dyslexia dataset is a modified version of the WMT14 English test set. This dataset was created to test the capabilities of SOTA machine translations models on dyslexic style text. This research was supported by AImpower.org.
In "Data/French_translated_data", each file within the dataset consists of a “.txt” or “.docx” file containing the translated sentences from AWS, Google, Azure and OpenAI.
In… See the full description on the dataset page:
https://huggingface.co/datasets/gpric024/wmt14_injected_synthetic_dyslexia.