This repository hosts SonoScene360, the real-world evaluation dataset introduced in SonoWorld: From One Image to a 3D Audio-Visual Scene (CVPR 2026). SonoWorld generates a 3D audio-visual scene with spatialized sound from a single input image.
[Paper] [Project website] [Official code]
The current dataset snapshot contains 10 scenes, 34 calibrated microphone poses, and 68 selected SN3D-normalized first-order Ambisonics (FOA)… See the full description on the dataset page:
https://huggingface.co/datasets/DerongJin/SonoScene360.