The first benchmark to systematically evaluate Reward Models' ability to assess long-term memory management in LLMs across contexts up to 128K tokens.
MemRewardBench is the first dedicated benchmark for evaluating Reward Models (RMs) in their ability to judge long-term memory management processes in Large Language Models. Unlike existing benchmarks that evaluate LLMs directly, MemRewardBench focuses on assessing how well RMs can evaluate… See the full description on the dataset page:
https://huggingface.co/datasets/LCM-Lab/MemRewardBench.