HaluMem: A Comprehensive Benchmark for Evaluating Hallucinations in Memory Systems
š Why We Define the HaluMem Evaluation Tasks
Limitations of Existing Frameworks
Most existing evaluation frameworks treat memory systems as black-box models, assessing performance only through end-to-end QA accuracy.
However, this approach has two major limitations:
It lacks a hallucination evaluation specifically designed for the characteristics of memory systems.⦠See the full description on the dataset page: https://huggingface.co/datasets/IAAR-Shanghai/HaluMem.