This model is a fine-tuned version of
google/mt5-large, trained for
Grammatical Error Correction (GEC) in
Thai for
L2 learners. It was developed as part of the research
"Grammatical Error Correction for L2 Learners of Thai Using Large Language Models", and represents the best-performing model in the study.
This model is based on the mT5-large architecture and was fine-tuned on the CTFL-GEC dataset, which contains human-annotated grammatical error corrections from L2 Thai learners. To improve generalization, the dataset was augmented using the Self-Instruct method with 200% additional synthetic pairs.
The model is capable of correcting sentence-level grammatical errors typical of L2 Thai writing, including issues with word order, omissions, and incorrect particles.
Evaluation was conducted on a held-out portion of the human-annotated dataset using common GEC metrics.