This repo contains 67,000 1-minute clips, amounting to approximately 1,117 hours for training, and 3,060 1-minute clips, amounting to roughly 51 hours for testing.
The dataset features an ontology of 170 sound classes and is generated by convolving sound event clips from FSD50K with simulated SRIRs (for training) or collected SRIRs from TAU-SRIR DB… See the full description on the dataset page:
https://huggingface.co/datasets/Jinbo-HU/PSELDNets.