VCog-Bench: Benchmarking Multimodal LLMs on Abstract Visual Reasoning
Description:
VCog-Bench is a publicly available zero-shot abstract visual reasoning (AVR) benchmark designed to evaluate Multimodal Large Language Models (MLLMs). This benchmark integrates two well-known AVR datasets from the AI community and includes a newly proposed MaRs-VQA dataset. The findings in VCog-Bench show that current state-of-the-art MLLMs and Vision-Language Models (VLMs), such as GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/vcog/vcog-bench.