We create MMEvalPro for more accurate and efficent evaluation for Large Multimodal Models. It is designed to avoid Type-I errors through a trilogy evaluation pipeline and more rigorous metrics. For each original question from existing benchmarks, human annotators augment it by creating one perception question and one knowledgeanchor question through a meticulous annotation process.
Data… See the full description on the dataset page: https://huggingface.co/datasets/MM-Diagnose/MMEvalPro.