The English Toxic Language Dataset for NLP is a synthetic dataset containing 10,000 labeled English comments designed for toxic language detection, content moderation, and natural language processing (NLP) research. It is suitable for developers, data scientists, and researchers building AI systems that can identify and classify toxic, abusive, or harmful text.
The dataset includes binary labels (Clean and Toxic), toxicity categories… See the full description on the dataset page:
https://huggingface.co/datasets/dheerubhadoria/English-Toxic-Language-Dataset-for-NLP.