This dataset contains 100 randomly selected video segments with comprehensive question-answer pairs for multimodal understanding tasks.
Action Identification: What actions are being performed
Attention Focus: What creates the overall mood and intensity
Attribute Transformation: How things change over time
Causal Reasoning: Why events happen and their causes… See the full description on the dataset page:
https://huggingface.co/datasets/ngqtrung/full-modality-sample-segments.