This repository contains VideoVista-CoTs, used in Uni-MoE-2.0 training.
This dataset samples a portion of data from LLaVA-Video-178K, SEED-Bench-R1, SR-91K, and STAR, and uses our automatic Video QA generation framework to perform multi-step reasoning annotations for filtered complex questions.
The automatic video QA generation codes and our VideoVista series are presented in VideoVista Family
If you find VideoVista-CulturalLingo useful for your… See the full description on the dataset page:
https://huggingface.co/datasets/HIT-TMG/VideoVista-CoTs.