AVRT-20K is a dataset of audio-visual reasoning traces generated through the AVRT (Audio-Visual Reasoning Transfer) pipeline. It provides structured chain-of-thought reasoning that explicitly integrates audio and visual evidence for answering multiple-choice questions about video content.
This dataset accompanies the paper:
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
Edson Araujo, Saurabhchand Bhati, M. Jehanzeb… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/AVRT-20K.