Model Description
We trained a distilled version of RoBERTa, using exerpts from three different emotion recognition datasets. The model was designed to classify texts
into categories based on Ekman's six basic emotions: anger, disgust, fear, joy, sadness, surprise, along with a neutral category.
Training Data
To even out language differences between data from different contexts, we selected three emotion recognition datasets from different domains.
This approach aims to mitigate language differences because of varying levels of noisiness across datasets.
From the domain of social media, we used GoEmotions (Demszky et al., 2020), which consists of approximately 54,000 online posts from the
platform Reddit. The dataset employed from the domain of dialogue datasets is the Multimodal Emotion Lines Dataset (MELD) (Poria et al., 2019).
MELD is a derived from the television series "Friends" and thus written by human authors. The dataset comprises approximately 13,000 utterances
in total. The final dataset is the International Survey on Emotion Antecedents and Reactions (ISEAR), which is a dataset for event-based data. The dataset was collected as part of a research project led by Scherer and Wallbott. During the project, students from
a range of academic disciplines and countries were invited to describe situations of experiencing one of seven distinct emotions, including joy, fear, sadness, disgust,
shame, and guilt (Swiss Center for Affective Sciences, n.D.). The dataset consists of approximately 7,000 examples of these descriptions.
We filtered the datasets as not all of them only include the six basic emotions. From the GoEmotions dataset, we only used data assigned to one emotion, filtering out
rows with multiple emotions. This process results in a total of 40,616 labelled examples.
Label Distribution:
| Label | No. of examples |
|---|
| Neutral | 22,457 |
| Joy | 4,454 |
| Anger | 3,968 |
| Sadness | 3,101 |
| Surprise | 2,538 |
| Disgust | 2,092 |
| Fear | 2,006 |
Dataset Distribution:
| Dataset | No. of examples |
|---|
| Train | 30,439 |
| Test | 5,868 |
| Validation | 4,309 |
Training Procedure
The pre-trained RoBERTa model is fine-tuned utilising the Trainer class from Hugging Face with a batch size of 64, a learning rate of 2e-6 and 15 epochs. After 14
epochs, the emergence of initial overfitting can be observed. However, only the best model is loaded at the end of the training process, which is the model from epoch
14.
Evaluation
The resulting model shows a test loss of 0.7696, which is relatively high, but not
unexpected, given the diverse data domains involved. The test accuracy is 0.7412,
while the F1 score is 0.7397.
Sources
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav
Nemade, and Sujith Ravi. GoEmotions: A Dataset of Fine-Grained Emotions.
In Proceedings of the 58th Annual Meeting of the Association for Computational
Linguistics, pages 4040–4054, 2020.
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik
Cambria, and Rada Mihalcea. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.
In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 527–536, 2019
Swiss Center for Affective Sciences. Research Material and Online Research. ht
tps://
www.unige.ch/cisa/research/materials-and-online-researc
h/research-material/, n.D. Last accessed: 16.04.2024.