BrazilianToxicTweetsClassification
An MTEB dataset
Massive Text Embedding Benchmark
ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from
a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity,
sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia,
Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/told-br.