This dataset is from our EMNLP'24 (main conference) paper MIBench: Evaluating Multimodal Large Language Models over Multiple Images
MIBench covers 13 sub-tasks in three typical multi-image scenarios: Multi-Image Instruction, Multimodal Knowledge-Seeking and Multimodal In-Context Learning.
Multi-Image Instruction: This scenario includes instructions for perception, comparison and reasoning across multiple input images. According to the… See the full description on the dataset page:
https://huggingface.co/datasets/StarBottle/MIBench.