Audio datasets referenced in Class-Agnostic Audio Repetition Counting. Contains two main sections:
RS, RSN and RVN: Synthetic datasets containing varying levels of noise. Uniformly 10 seconds long, with 0-8 repetition events contained in each sample.
Clocks, Heartbeats and Dolphins: Real-world derived samples with variable length across mechanical, ecological and medical domains.
Relevant code can be found in this repo.
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/Hamozwa/RepeatAudio.