VideoChat3-Academic2M is the academic video instruction data used by VideoChat3. It re-annotates public academic video datasets for video captioning, video question answering, and fine-grained motion understanding.
The dataset follows an evidence-grounded annotation enhancement pipeline. Short answers, option-only labels, and concise captions are rewritten into richer instruction-following responses that mention visible objects, actions, scenes, temporal… See the full description on the dataset page:
https://huggingface.co/datasets/MCG-NJU/VideoChat3-Academic2M.