This dataset is designed for training models to detect incoherence in text. It includes various types of incoherence, such as grammatical errors, word soup, random words, and run-on sentences.
Dataset Details
Languages: English, Spanish, French, German, Chinese, Japanese, Russian, Arabic, Hindi
Size: ~27,000 samples
Types of Incoherence: Grammatical errors, word soup, random words, run-ons, random tokens, random bytes.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/SuccubusBot/incoherent-text-dataset.