Audio-visual spatial-audio dataset built by placing sound-emitting objects in
17 Replica indoor scenes with a Unity/PhysX physical placement stage, then
rendering binaural room impulse responses (RIR) with
SoundSpaces 2.0 (habitat-sim
AudioSensor, per-class acoustic materials). Each case ships the listener's
first-person render, per-source binaural RIRs, dry source audio, the
RIR-convolved per-source audio, a mixed binaural… See the full description on the dataset page:
https://huggingface.co/datasets/mumbumble/GAZE_soundspace.