This dataset (MSRVTT_PseudoImageCaptions.json) is released as part of our CVPR 2025 paper:
DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
ArXiv:
https://arxiv.org/abs/2506.08887
Github:
https://github.com/LunarShen/DsicoVLA
Citation
@inproceedings{shen2025discovla,
title={DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval},
author={Shen, Leqi and Gong, Guoqiang and… See the full description on the dataset page:
https://huggingface.co/datasets/LeqiShen/DiscoVLA.