This dataset contains audio-text pairs in the webdataset format.
The audio files are short speech segments from publicly available videos & the texts are descriptions of emotions the speakers seems to be feeling. Some captions also describe the speakers gender and age.
All files with the substring "part1" in the name contain unique audio files with unique captions.
All files with the substring "part2" , "part3", ... in the name contain the same audio files as in "part1", but with different… See the full description on the dataset page:
https://huggingface.co/datasets/EQ4You/Emotional_Speech.