MovieCORE is a comprehensive video question answering (VQA) dataset specifically designed to evaluate and probe deeper cognitive understanding of movie content. Unlike traditional VQA datasets that focus on surface-level visual understanding, MovieCORE challenges models to demonstrate sophisticated reasoning about narrative structures, character development, thematic elements, and complex temporal relationships within cinematic content.
🗂️ Data Preparation
The MovieCORE dataset builds upon video content from MovieChat. To get started:
Video Data
Download the video files from MovieChat's HuggingFace repositories:
We have provided the script to run HERMES (ICCV'25) on MovieCORE. Please check out the linked project.
📊 Evaluation Dimensions
MovieCORE employs a comprehensive multi-dimensional evaluation framework to assess model performance across different aspects of cognitive understanding:
Dimension
Description
🎯 Accuracy
Measures semantic similarity between predicted and ground truth answers
📋 Comprehensiveness
Assesses coverage of all key aspects mentioned in the ground truth
🧠 Depth
Evaluates level of reasoning and insight demonstrated in predictions
🔍 Evidence
Checks quality and relevance of supporting evidence provided
🔗 Coherence
Measures logical flow, organization, and clarity of responses
Each dimension provides unique insights into different cognitive capabilities required for deep video understanding.
💻 Usage
Evaluation Script
Evaluate your model's performance on MovieCORE using our evaluation script:
Your predictions should follow this JSON structure:
json
1{2"video_1.mp4":[3{4"question":"How does the video depict the unique adaptations of the species in the Sahara Desert, and what roles do these species play in their ecosystem?",5"answer":"The ground truth answer.",6"pred":"Your model's prediction.",7"classification":"the question classification"8},9{10"question":"The second question for video 1?",11"answer":"The ground truth answer.",12"pred":"Your model's prediction.",13"classification":"the question classification"14}15],16"video_2.mp4":[17{18"question":"The only question for video 2",19"answer":"The ground truth answer.",20"pred":"Your model's prediction.",21"classification":"the question classification"22}23]24}
📈 Output
The evaluation script provides:
Overall scores across all dimensions
Classification-specific performance metrics
Detailed breakdowns for comprehensive analysis
📚 Citation
If you use MovieCORE in your research, please cite our paper:
bibtex
1@misc{faure2025moviecorecognitivereasoningmovies,
2 title={MovieCORE: COgnitive REasoning in Movies},
3 author={Gueter Josmy Faure and Min-Hung Chen and Jia-Fong Yeh and Ying Cheng and Hung-Ting Su and Yung-Hao Tang and Shang-Hong Lai and Winston H. Hsu},
4 year={2025},
5 eprint={2508.19026},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2508.19026},
9}
🤝 Contributing
We welcome contributions to MovieCORE! Please feel free to:
Report issues or bugs
Suggest improvements or new features
Submit baseline implementations
Provide feedback on the evaluation framework
📄 License
This dataset is provided under the MIT License. See LICENSE for more details.
🎬 Advancing Video Understanding Through Cognitive Evaluation 🎬