SpotSound-Bench: A 'Needle-in-a-Haystack' Evaluation for Audio Temporal Grounding
Benchmark Summary
SpotSound-Bench is a challenging temporal grounding benchmark designed to evaluate Large Audio-Language Models (ALMs).
Existing benchmarks for audio temporal grounding often feature high ratios of target-window duration to full audio clip duration, which fail to simulate real-world scenarios where short events are obscured by dense background sounds. To bridge… See the full description on the dataset page: https://huggingface.co/datasets/Loie/SpotSound-Bench.