Dataset Card for T2AV-Compass
Dataset Details
Dataset Description
T2AV-Compass is a unified benchmark for evaluating Text-to-Audio-Video (T2AV) generation, targeting not only unimodal quality (video/audio) but also cross-modal alignment & synchronization, complex instruction following, and perceptual realism grounded in physical/common-sense constraints.
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/T2AV-Compass.