Based on p1atdev/open2ch
We apply keyword-based filtering to collect toxic texts
We use Perspective API to filter non-toxic texts from the original corpus
3k texts for each class, toxic (label=1) and non-toxic (label=0) texts
perspective_api_score is a prediction of toxicity score by the Perspective API