SeLex-RT Outputs: Multilingual Round-Trip Synthetic GEC Data
Dataset Summary
This dataset contains synthetic training data for low-resource Grammatical Error Correction (GEC), generated via a round-trip machine translation (RT) pipeline inspired by SeLex-RT from Low-Resource Grammatical Error Correction: Selective Data
Augmentation with Round-Trip Machine Translation (Gomez & Rozovskaya, 2025). The pipeline targets lexical errors — the class of errors most systematically… See the full description on the dataset page: https://huggingface.co/datasets/gabelev/selex-rt-outputs.