Do Audio-Visual Large Language Models Really See and Hear?
Paper | Project Page | GitHub
This dataset is part of the first mechanistic interpretability study of Audio-Visual Large Language Models (AVLLMs). It is designed to analyze how audio and visual features evolve and fuse through different layers of models like Qwen 2.5 Omni. The data specifically supports investigating modality bias and how models handle conflicting information between audio and vision.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/gamma-lab-umd/counterfactual-av-eval.