VGGSynth1: Synthetic Audio-Visual Dataset (Part 2)
How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization? (CVPR 2026 Highlight)
Authors
VGGSynth2 is a high-fidelity synthetic clone of the VGGSound dataset, built using state-of-the-art generative models.
This dataset is designed to explore the boundaries and utility of synthetic… See the full description on the dataset page: https://huggingface.co/datasets/swimmiing/VGGSynth2.