A dataset containing English speech with grammatical errors, along with the corresponding transcriptions. Utterances are synthesized using a
text-to-speech model, whereas the grammatically incorrect texts come from the C4_200M synthetic dataset.
The Synthesized English Speech with Grammatical Errors (SESGE) dataset was developed to support the DeMINT project
developed at Universitat… See the full description on the dataset page:
https://huggingface.co/datasets/Transducens/sesge.