BEWO-1M: Open Source Spatial Audio Dataset
Introduction
To better facilitate the advancement of multimodal guided spatial audio generation models, we have developed a dual-channel audio dataset named Both Ears Wide Open 1M (BEWO-1M) through rigorous simulations and GPT-assisted caption transformation.
Totally, we constructed 2.8k hours of training audio with more than 1M audio-text pairs and approximately 17 hours of validation data with 6.2k pairs.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/spw2000/BEWO-1M.