BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla
Dataset Overview
Hate detection (binary) and target identification (multi-label) dataset on 37.3 k transliterated Bangla samples.
~38% hate samples and 7 target classes: Political, Religious, Gender, Personal Offense, Abusive/Violence, Origin, and Body Shaming.
Sourced from 26 YouTube channels of 3 categories: News & Politics, People & Blogs, and Entertainment.
Each sample is labeled by 3… See the full description on the dataset page: https://huggingface.co/datasets/aplycaebous/BanTH.