RMHackBench is a mechanism-oriented benchmark designed to systematically analyze and diagnose reward hacking behaviors in Multimodal Reward Models (MRMs).
Unlike standard Multimodal Large Language Model (MLLM) benchmarks that evaluate general reasoning capabilities, RMHackBench specifically tests whether a reward model is evaluating the genuine visual grounding of an answer or merely being misled by superficial… See the full description on the dataset page:
https://huggingface.co/datasets/sdfss1/RMHackBench.