This dataset is a refined version of the PleIAs/ToxicCommons collection, focusing on historical texts labeled for content that may be considered objectionable by modern standards (what the authors of the dataset deem "toxic").
The cleaned dataset contains 1 051 027 rows, each representing a text sample with associated toxicity scores across five dimensions:
Race and origin-based bias
Gender and sexuality-based bias
Religious bias
Ability bias
Violence and abuse… See the full description on the dataset page:
https://huggingface.co/datasets/agentlans/PleIAs-ToxicCommons.