Authors: Yongkang Du, Xiaohan Zou, Minhao Cheng, Lu Lin · Pennsylvania State University
CARV evaluates whether multimodal LLMs can compose transformation rules from multiple image pairs via logical set operations. Given n context pairs each depicting an atomic visual change, the model must synthesize a new rule through Union (∪), Intersection (∩), or… See the full description on the dataset page:
https://huggingface.co/datasets/duyongka/CARV.