This dataset is produced from using Detoxify (
https://github.com/unitaryai/detoxify) on the dataset:
content not marked for toxicity
content marked for toxicity incorrectly
some content marked with high scores that doesn't seem toxic
some content not marked when clearly offensive
However, the bulk seems to be fairly right on the mark, so I'm… See the full description on the dataset page:
https://huggingface.co/datasets/jtatman/pippa_deduped_detoxify_score.