This dataset was developed as part of a project for CS5246, taken in AY23/24 S2 at the National University of Singapore. My group and I intially processed the data from Reddit,
turned it into a binary classification task, and addressed dataset imbalance by sourcing extra samples from a publicly available dataset here.
The training set consists of 10,872 samples,
which are balanced between 5,134 samples labelled as
'not morally wrong' and 5,738… See the full description on the dataset page:
https://huggingface.co/datasets/jeanong2/AITA-datasets.