This is an abusive/offensive language detection dataset for Albanian. The data is formatted
following the OffensEval convention, with three tasks:
-
Subtask A: Offensive (OFF) or not (NOT)
-
Subtask B: Untargeted (UNT) or targeted insult (TIN)
-
Subtask C: Type of target: individual (IND), group (GRP), or other (OTH)
-
The subtask A field should always be filled.
-
The subtask B field should only be filled if there's "offensive" (OFF) in A.
-
The subtask C field should only be filled if there's "targeted" (TIN) in B.
The dataset name is a backronym, also standing for "Spoken Hate in the Albanian Jargon"