[π Paper] [π Dataset ] [β¨ Github]
@article{li2024videovista,
title={Videovista: A versatile benchmark for video understanding and reasoning},
author={Li, Yunxin and Chen, Xinyu and Hu, Baotian and Wang, Longyue and Shi, Haoyuan and Zhang, Min},
journal={arXiv preprint arXiv:2406.11303},
year={2024}
}
VideoVista-Train consists of 114,581 training samples derived from 3,838 video clips.
These samples cover⦠See the full description on the dataset page:
https://huggingface.co/datasets/Uni-MoE/VideoVista_Train.