OffensEval 2020 features a multilingual dataset with five languages. The languages included in OffensEval 2020 are:
The annotation follows the hierarchical tagset proposed in the Offensive Language Identification Dataset (OLID) and used in OffensEval 2019.
In this taxonomy we break down offensive content into the following three sub-tasks taking the type and target of offensive content into account.
The following sub-tasks were organized:
The English training data isn't included here (the text isn't available and needs rehydration of 9 million tweets;
see
https://zenodo.org/record/3950379#.XxZ-aFVKipp)