This dataset is designed to train language models to identify and correct a wide range of grammatical errors and stylistic issues in Japanese text.
The data consists of pairs of incorrect and correct sentences, along with metadata that classifies the type of error and provides additional context.
The dataset was created by both manual curation from discussions in Japanese learning communities and synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/huytd189/japanese-grammar-correction.