Common-O, inspired by cognitive tests for humans, probes multimodal LLMs' ability to reason across scenes by asking "what’s in common?"
We have two subsets: Common-O (3 - 8 objects) and Common-O Complex (8 - 16 objects).
Multimodal LLMs excel at single image perception, but struggle with multi-scene reasoning
Evaluating a Multimodal LLM on Common-O
import… See the full description on the dataset page:
https://huggingface.co/datasets/facebook/Common-O.