This dataset contains preprocessed data from the VGGSound dataset, specifically processed using the VGGSound-AVEL50k subset for cross-modal knowledge distillation research. The preprocessing is optimized for MST-Distill (Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation) method.
This preprocessing work is based on the VGGSound-AVEL50k subset from: jasongief/CPSP: [2023 TPAMI] Contrastive Positive Sample Propagation along the… See the full description on the dataset page:
https://huggingface.co/datasets/Gray1y/VGGSound-50k.