The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is formed by 10000 thousand unpublished annotated tweets collected in 2023.
This repository is organized as follows:
root.
├── annotations : classification given by annotators
├── raw corpus : dataset before… See the full description on the dataset page:
https://huggingface.co/datasets/victoriadreis/TuPY_dataset_multilabel.