The Offensice Language Identification Dataset (OLID) contains 14,100 annotated tweets from Twitter, annotated with three subcategories via crowdsourcing and has been released together with
the paper Predicting the Type and Target of Offensive Posts in Social Media.
Previous datasets mainly focused on detecting specific types of offensive messages (hate speech, cyberbulling, etc.) but did not consider offensive language as a whole.
This dataset is… See the full description on the dataset page:
https://huggingface.co/datasets/christophsonntag/OLID.