This is the repository for the S3D dataset published at EMNLP 2022. The dataset can help build sarcasm detection models.
The S3D dataset is our silver standard dataset of 100,000 tweets labelled for sarcasm using weak supervision by our BERTweet-sarcasm-combined model.
These tweets can be accessed by using the Twitter API so that they can be used for other experiments.
S3D contains 38879… See the full description on the dataset page:
https://huggingface.co/datasets/surrey-nlp/S3D-v1.