The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML)
and natural language processing (NLP) techniques. TuPy is comprised of 10,000 (ten thousand) unpublished, annotated, and anonymized documents collected
on Twitter (currently known as X) in 2023.
This repository is organized as follows:
root.
├── binary : binary… See the full description on the dataset page:
https://huggingface.co/datasets/Silly-Machine/TuPy-Dataset.