RM-Bench style dataset built from PopQA for evaluating factuality bias in reward models.
Same 6-surface × 3-style structure as GSM8K-RMBench but for factual QA instead of math reasoning.
Existing RM style-bias benchmarks (e.g., GSM8K-RMBench, RM-Bench) focus on mathematical/reasoning correctness. This dataset extends the same paradigm to factual knowledge recall — measuring whether reward models can distinguish factually correct answers… See the full description on the dataset page:
https://huggingface.co/datasets/xxccho/popqa-rmbench.