AudioSet[1] consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube.
Some clips are missing on YouTube, so the number of files downloaded is different from time to time.
This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set.
The pre-process script can be found at… See the full description on the dataset page:
https://huggingface.co/datasets/yangwang825/audioset.