This is the synthetic training corpus for Urdu Grammatical Error Correction (GEC) introduced in the paper "Corpora Generation for Urdu Grammatical Error Correction" (Accepted at ACL Findings).
The dataset contains approximately 1.27 million sentence pairs. It was created by mining naturally occurring error patterns from Urdu Wikipedia revision histories and re-inflicting them onto clean text (Makhzan Corpus) using a kernel-based… See the full description on the dataset page: https://huggingface.co/datasets/Kagura-Ahad-123/UrduGEC-Synthetic.