This repository contains the YT-NTU-AVQ dataset associated with the ICASSP 2026 paper:
Scaling Audio-Visual Quality Assessment Dataset via Crowdsourcing
YT-NTU-AVQ is a audio-visual quality assessment dataset consisting of 1,620 audio-visual sequences collected from YouTube, where 80% of the videos are sourced from AudioSet.
All videos are released under Creative Commons Attribution 3.0 (CC BY 3.0, YouTube Creative Commons).
For more details, please refer to the… See the full description on the dataset page:
https://huggingface.co/datasets/ntu-avqa/YT-NTU-AVQ.