A long-context urban-audio QA benchmark across 10 task types and
4 context lengths (100 s, 15 min, 30 min, 1 h). Each soundscape is
synthesised by Scaper from
UrbanSound8K foreground events over TUT acoustic-scene backgrounds, sampled
at 16 kHz mono PCM_32. Two of the ten tasks
(anomaly_detection, anomaly_localization) draw from a parallel pool
where every soundscape contains exactly one out-of-vocabulary event from
ESC-50 (glass_breaking or crying_baby).
This… See the full description on the dataset page:
https://huggingface.co/datasets/nz00shuuuu/urbansound-haystack.