Benchmark type:
SEED-Bench-2 is a comprehensive large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs), featuring 24K multiple-choice questions with precise human annotations.
It spans 27 evaluation dimensions, assessing both text and image generation.
Benchmark date:
SEED-Bench was collected in November 2023.
Paper or resources for more information:
https://github.com/AILab-CVC/SEED-Bench
License:… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Bench-2.