MR²Bench (Multi-Response Multimodal Reward Bench) is a pair of benchmarks for evaluating multimodal reward models on N-way ranking tasks. Unlike existing benchmarks that only support pairwise comparisons, MR²Bench provides N-way human-annotated rankings over responses from multiple diverse models, enabling evaluation of both pairwise and listwise ranking capabilities.
This benchmark is described in our paper:
You Only Judge Once: Multi-response Reward Modeling in a… See the full description on the dataset page:
https://huggingface.co/datasets/yinuoy/MR2Bench.