Chunk AudioSet every 0.25 seconds and 0.5 seconds and predict using MIT/ast-finetuned-audioset-10-10-0.4593, after that only filter and labels predicted overlapped with the gold labels.
Only chunk on balanced_train_segments.zip and eval_segments.zip.
0 Speech : 97017
1 Male speech, man speaking : 1508
2 Female speech, woman speaking : 153
3 Child speech, kid speaking : 485
4 Conversation : 0
5 Narration, monologue… See the full description on the dataset page:
https://huggingface.co/datasets/mesolitica/AudioSet-Chunk.