DAVE is a diagnostic benchmark for evaluating audio-visual models, ensuring both modalities are required and providing fine-grained error analysis to reveal specific failures. Researchers can use DAVE to test and compare audio-visual models, refine multi-modal architectures, or develop new methods for audio-visual alignment. It is not intended for training large-scale models but for targeted evaluation… See the full description on the dataset page:
https://huggingface.co/datasets/gorjanradevski/dave.