This dataset is a set of samples for training and testing the spell checking, grammar error correction and ungrammatical text detection models.
The dataset contains two splits:
test.json contains samples hand-selected to evaluate the quality of models.
train.json contains synthetic samples generated in various ways.
The purpose of creating the dataset was to test an internal spellchecker for a generative poetry project, but it can also be useful in other projects… See the full description on the dataset page:
https://huggingface.co/datasets/inkoziev/spellchecker.