SO-Dataset is a large-scale spatial audio dataset in first-order ambisonics (FOA) format. Each example contains FOA waveform and spatial event annotations in DCASE-style CSV files. The dataset combines simulated spatial scenes and real FOA recordings, and all sound event labels are mapped into a unified 63-class sound event taxonomy based on the FSD50k dataset.
The public release stores audio and annotations as tar shards. The tar files… See the full description on the dataset page:
https://huggingface.co/datasets/dieKarotte/SO-Dataset.