The {M}arathi {O}ffensive {L}anguage {D}ataset (MOLD) contains a collection of 2500 annotated Marathi tweets.
The files included are:
MOLD
│ README.md
└───data
│ MOLD_train.csv
│ MOLD_test.csv
MOLD_train.csv: contains 1,875 annotated tweets for the training set.
MOLD_test.csv: contains 625 annotated tweets for the test set.
The dataset was annotated using crowdsourcing. The gold labels were assigned taking the… See the full description on the dataset page:
https://huggingface.co/datasets/tharindu/MOLD.