This dataset is used for Bangla GEC. The dataset is divided into test and train part for all the categories of data.
If you are using this dataset please cite:
@inproceedings{bhattacharyya-bhattacharya-2025-leveraging,
title = "Leveraging {LLM}s for {B}angla Grammar Error Correction: Error Categorization, Synthetic Data, and Model Evaluation",
author = "Bhattacharyya, Pramit and
Bhattacharya, Arnab",
editor = "Che, Wanxiang and
Nabende, Joyce and… See the full description on the dataset page:
https://huggingface.co/datasets/Vacaspati/Vaiyakarana.