LRSA-XLMR-Sentiment-ID
Model Description
LRSA-XLMR-Sentiment-ID is a fine-tuned XLM-RoBERTa (XLM-R) model for Indonesian sentiment analysis developed as part of the study:
Stress-Testing Large Language Models (LLMs) against Code-Mixing and Distributional Shifts in Low-Resource NLP
The model was trained and evaluated on a large Indonesian sentiment corpus containing approximately 75,000 samples collected from social media and news domains.
Task
Sentiment Classification
Labels:
- Negative
- Neutral
- Positive
Dataset
Sources:
- YouTube comments
- Indonesian news articles and comments
Domains:
- Politics
- Economics
- Social issues
- Public policy
Statistical Significance
McNemar testing demonstrated statistically significant differences between XLM-R and the baseline SVM model (p < 0.001), confirming the effectiveness of transformer-based architectures under low-resource Indonesian sentiment classification settings.
Repository
Code, experimental results, figures, and supplementary materials:
Authors
Zia Ul Rehman Zafar
Dedi Gunawan
Endang Wahyu Pamungkas
Widi Widayat
Helmi Imaduddin
Department of Informatics Engineering
Universitas Muhammadiyah Surakarta, Indonesia
Citation
If you use this model, please cite:
Zafar, Z. U. R., Gunawan, D., Pamungkas, E. W., Widayat, W., & Imaduddin, H.
Stress-Testing Large Language Models (LLMs) against Code-Mixing and Distributional Shifts in Low-Resource NLP.
License
MIT License