Views
No views yet
1from transformers import AutoTokenizer, AutoModel
2tokenizer = AutoTokenizer.from_pretrained("indolem/indobertweet-base-uncased")
3model = AutoModel.from_pretrained("indolem/indobertweet-base-uncased")| Models | Sentiment | Emotion | Hate Speech | NER | Average | |||
|---|---|---|---|---|---|---|---|---|
| IndoLEM | SmSA | EmoT | HS1 | HS2 | Formal | Informal | ||
| mBERT | 76.6 | 84.7 | 67.5 | 85.1 | 75.1 | 85.2 | 83.2 | 79.6 |
| malayBERT | 82.0 | 84.1 | 74.2 | 85.0 | 81.9 | 81.9 | 81.3 | 81.5 |
| IndoBERT (Willie, et al., 2020) | 84.1 | 88.7 | 73.3 | 86.8 | 80.4 | 86.3 | 84.3 | 83.4 |
| IndoBERT (Koto, et al., 2020) | 84.1 | 87.9 | 71.0 | 86.4 | 79.3 | 88.0 | 86.9 | 83.4 |
| IndoBERTweet (1M steps from scratch) | 86.2 | 90.4 | 76.0 | 88.8 | 87.5 | 88.1 | 85.4 | 86.1 |
| IndoBERT + Voc adaptation + 200k steps | 86.6 | 92.7 | 79.0 | 88.4 | 84.0 | 87.7 | 86.9 | 86.5 |
1@inproceedings{koto2021indobertweet,
2 title={IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization},
3 author={Fajri Koto and Jey Han Lau and Timothy Baldwin},
4 booktitle={Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021)},
5 year={2021}
6}