VideoChat3-OL617K is the online video instruction data used by VideoChat3. It is designed to train proactive streaming video assistants that continuously observe incoming video, accumulate visual evidence, and respond at the appropriate moment.
The dataset converts video-question-answer triples into causal streaming supervision. Visual clue intervals are first localized and verified, then transformed into streaming sequences with explicit response-state tokens:… See the full description on the dataset page:
https://huggingface.co/datasets/MCG-NJU/VideoChat3-OL617k.