Jopara (Guarani-dominant mixed with Spanish) Sentiment Analysis (JOSA) corpus.
We collected a dataset of tweets primarily written in Guarani (and Jopara, a code-switching language that combines Guarani and Spanish) and annotated them polarity (positive, neutral and negative classes) in sentiment analysis. We created two corpora:
Unbalanced with 3941 tweets, and
Balanced with 1526 tweets, created from Unbalanced set.
The statistics for the Jopara… See the full description on the dataset page:
https://huggingface.co/datasets/mmaguero/gn-jopara-sentiment-analysis.