MPBench, a comprehensive benchmark for assessing the effectiveness of multimodal process reward models (PRMs) in various scenarios, achieved through three evaluation paradigms: Step Correctness, Answer Aggregation, and Reasoning Process Search.
More details on loading and using the data are at our github page.
If you do find our code helpful or use our benchmark dataset, please citing our… See the full description on the dataset page:
https://huggingface.co/datasets/xuzhaopan/MPBench.