WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset).
avg. audio duration (s)avg. text length
FreeSound… See the full description on the dataset page:
https://huggingface.co/datasets/cvssp/WavCaps.