Clone from "friedrichor/MSR-VTT".
MSRVTT contains 10K video clips and 200K captions.
We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field.
Train:
@inproceedings{xu2016msrvtt,
title={Msr-vtt: A large video description dataset… See the full description on the dataset page:
https://huggingface.co/datasets/VLM2Vec/MSR-VTT.