AudioSet data: {ytb_id}
{start_second*1000}{end_second
1000}
AudioSet Strongly Labeled Subset: {ytb_id}_{start_second1000} (10s clip from the the start second)
VggSound: {ytb_id}_{start_second} (10s… See the full description on the dataset page:
https://huggingface.co/datasets/OpenSound/EzAudioCaps.