The toxic-data-bg dataset consists of 4,384 manually annotated sentences across four categories: toxic language, medical terminology, non-toxic language, and terms related to minority communities.
The dataset is an extention of the "Hate speech detection in Bulgarian" dataset and consist of sentences of:
BG-Jargon
BG-Nationalisti forum
BG-Mamma forum
Proud.bg forum… See the full description on the dataset page:
https://huggingface.co/datasets/sofia-uni/toxic-data-bg.