This dataset is built upon Shunian/kaggle-mbti-cleaned to address the sample imbalance problem.
Thanks to the Parrot Paraphraser and NLP AUG, some of the skewness issue are addressed in the training data, make it grows from 328,660 samples to 478,389 samples in total.
View GitHub for more information