This dataset is curated from internet videos to support research in dog vocalization detection using both weak and strong supervision.
It contains approximately 7,500 seconds of strongly labeled training audio
Over 9,000 seconds of weakly labeled clips sourced from AudioSet are included.
The dataset also provides 24 hours of unlabeled audio clips from our own collection.
To simulate realistic conditions, some clips feature dogs present without barking… See the full description on the dataset page:
https://huggingface.co/datasets/ArlingtonCL2/Barkopedia-Dog-Vocal-Detection.