CoreCognition: A Core Knowledge Benchmark for Multi-modal Large Language Models
Dataset Description
CoreCognition is a large-scale benchmark encompassing 12 core knowledge grounded in developmental cognitive science, designed to evaluate the fundamental core abilities of Multi-modal Large Language Models (MLLMs).
While MLLMs demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks… See the full description on the dataset page: https://huggingface.co/datasets/williamium/CoreCognition.